Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Knowledge bases”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Artificial Reasoning System for Symptom-Based Conditional Failure Probability Estimation Using Bayesian Network

Advances in nuclear power technologies require enhanced capabilities for operator advice and autonomous control. One of the first tasks in the development of such capabilities is the formulation of symptom-based conditional failure probabilities for structures, systems, and components (SSCs) of interest, for which the primary goal is to aid plant personnel in deducing the probabilistic performance status of the monitored SSCs and in detecting impending faults/failure. The task of conditional failure probability estimation is a bidirectional inference problem and shall be logically tackled by the Bayesian network (BN) approach. As a knowledge-based artificial intelligence tool and a probabilistic graphical model, BN offers the capability of reasoning under uncertainty and graphical representation emulating the physical behavior of the target SSC. This paper provides a systematic overview of the BN technique and the software tools for handling implementation of BN models, along with the associated knowledge representation and reasoning paradigm. Both operational data and expert judgement can be readily incorporated into the knowledge base of a BN model. The challenges with data availability are highlighted, and the general approach to target SSC identification is presented. Our focus is upon failure-prone and risk-important balance of plant assets, especially cases having strong operator involvement. An exemplary case study on the failure of a motor-driven centrifugal pump is also conducted to demonstrate the usefulness and technical feasibility of the proposed artificial reasoning system using an expert system shell.

Zhao, Xingang↗

A High-Throughput Computing Infrastructure to Generate Custom, Open Community Geothermal Datasets

The most significant challenge facing geothermal research, development, and deployment is a lack of comprehensive datasets describing the geological and economical properties of North America. Automated knowledge base construction, the process of designing algorithms to analyze text and images to programmatically build new datasets, is one possible solution to this problem. The xDD library of full-text scientific articles (https://xdd.wisc.edu) is one of the largest collections of open and controlled-access scientific documents available for knowledge base construction in the world, but it has been underutilized by experts in geothermal research. The xDD development team attributed the lack of engagement by software developers and geothermal researchers to two perceived shortcomings of the system. First, the workflow for obtaining data from xDD for local development and testing of data mining applications was unnecessarily abstruse and required significant manual intervention by xDD systems administrators. Second, although xDD already held articles from a broad cross-section of scientific literature with an emphasis on the geosciences, it did not have an explicit set of geothermal research documents that could serve as the nucleus of a geothermal data mining application. To address these issues, the Automated Data Extraction PlaTform (ADEPT) was proposed to extend the data distribution capabilities of the xDD system. The ADEPT extension added the following four key features to xDD: 1) integration of National Geothermal Data System (NGDS) documents into the xDD library to provide an explicitly geothermally-themed collection; 2) improved RESTful (i.e., https-protocol driven) web services for external partners to access xDD data for machine learning application development; 3) a web platform for end-users and xDD administrators to coordinate the development of data mining applications from the initial step of browsing available documents to the final stage of deploying a production-quality machine learning application on high-throughput computing infrastructure; and 4) the development of demonstration data mining applications to illustrate the new workflow to potential collaborators. A total of 21,674 geothermal documents from NGDS were fully ingested into the xDD library and the associated metadata is publicly available through the xDD web services; furthermore, the ADEPT web platform is now publicly accessible and fully live at https://xdd.wisc.edu/adept/.

15 GEOTHERMAL ENERGY↗

CLARIFYING THE NEXUS BETWEEN LIFE CYCLE ASSESSMENT AND CIRCULARITY INDICATORS: A SETAC/ACLCA INTEREST GROUP

Purpose Improving the circularity of resources is important to the sustainability of consumer goods. Current research has indicated that circularity practices and circular economy (CE) methods do not always reduce environmental impacts. The aim of this research is to investigate the adoption of the life cycle assessment (LCA) methodology to improve the environmental impacts of circularity practices. Methods As part of the Society for Environmental Toxicology And Chemistry (SETAC) forum, an interest group (IG) on Circularity and LCA was formed in partnership with the American Center for Life Cycle Assessment (ACLCA) to tackle methodological and technical issues related to circularity in LCA. The IG’s research approach is summarized in four key steps: defining goals and objectives, literature review and gap analysis, ideation, and experimentation. The twelve active persons within this IG meet monthly and have been divided into four sub-working groups (sub-WGs) so that complementary tasks can be completed concurrently in an effective manner. Each sub-WG meets monthly and reports back to the main group for collaboration and brainstorming to meet the research objectives. Results and discussion First, the sub-WG #1, focusing on the “pool of circularity and LCA-based indicators”, analyzed the complementarity between two of the most used circularity indicators and LCA. Second, the sub-WG #2, working on the “evaluation of CE loops performance through LCA”, built a mind map of pain points that reflect the challenges that the LCA practitioners face when combining LCA with CE approaches. Third, the sub-WG #3, dealing with the “trade-offs between circularity and sustainability”, highlighted key alignments and/or conflicts between circularity and sustainability performance depending on the scope, product, industry, or system of analysis. Fourth, the sub-WG #4, focusing on “business and industrial cases”, plans to leverage the knowledge base developed within this IG to develop use cases documenting the benefits and challenges associated with CE-related loops modeling in LCA. Conclusions The first findings of this SETAC/ACLCA IG provide a state-of-the-art overview of the synergists of LCA methodology and the CE measurement frameworks reported in the literature. To move forward and capitalize on the first findings, one valuable point will be to discuss and provide concrete solutions to the pain points that emerged when considering circularity in LCA. Eventually, the knowledge base and resources created within this IG ultimately aim to support the proper application of LCA for practitioners in CE contexts, and could provide relevant inputs for the ISO Technical Committee ISO/TC 323 working on the upcoming standard for the measurement of CE performance.

Life cycle assessment, circular economy, circulari↗

Enhancing Oxygen Stability in Low-Cobalt Layered Oxide Cathode Materials by Three-Dimensional Targeted Doping

In this project, we propose to develop a new concept and a generic platform that can lead to the greatly enhanced stabilization of all high-energy cathode materials, and in particular high-nickel (Ni) and low-cobalt (Co) oxides. The new concept is a 3D doping technology that hierarchically combines surface and bulk doping. We will use surface doping to stabilize the surface of primary particles and also introduce dopants in the bulk to further enhance oxygen stability, conductivity, and structural stability in low-Co oxides under high voltage and deep discharging operating conditions. This new concept not only will deliver a low-cost, high-energy cathode but also will provide a generic method that can stabilize all high-energy cathodes. The proposed novel 3D doping approach is poised to resolve some longstanding challenges in fundamental doping effects on battery materials as well as to reduce Li-ion batteries’ cost and improve their safety, energy density, and lifetime. To tackle this problem, we have formed a highly complementary multi-university/national labs/industry team to enable a doping-central and systematic investigation of low-Co materials and create a knowledge base for many electrode materials to be used in advanced electric vehicles. The successful execution of the proposed project relies on five components that can be carried out by the complementary team members: (1) a theoretical investigation of the surface and bulk stabilizing dopants (Persson), (2) a precise synthesis of materials with targeted doping (Lin and Xin), (3) development of electrolytes for high-Ni low-Co oxides (Xu), (4) multi-scale characterization of the structures and their interfaces by scanning transmission electron microscopy (SEM) and synchrotron X-ray imaging and spectroscopy tools (Xin and Lin), and (5) pouch cell-level integration (Fan). The UCI-led project will enable a doping-central and systematic investigation of low-Co materials and create a knowledge base for many electrode materials to be used in advanced electric vehicles.

25 ENERGY STORAGE↗

Leveraging Large Language Models for Real-World Data Evidence: A Framework for Automated Treatment Extraction and Data Harmonization

Background: The ability to comprehensively collect treatment information from cancer patient medical records would enable studies to evaluate real-world benefits and risks tied to specific treatments. Currently, it is difficult to system- atically collect high-quality treatment information because it is often stored in unstructured text. Manually extracting and standardizing drug and regimen data is time-intensive. Recent advances in large language models (LLMs) offer a potential solution for automated extraction of structured treatment information from clinical text. Objective: This study systematically evaluates the utility of four LLMs from the Llama family for automated extraction of oncology treatment information from clinical text. This information can guide researchers using cancer registry data to provide insights into cancer care and outcomes beyond clinical trials. Methods: Four instruction-tuned Llama models with varying parameter counts (1B, 3B, 8B, and 70B) were evaluated for their ability to extract treatment information from clinical documents. A unified oncology knowledge base integrating seven major public data sources was developed to standardize and normalize extracted entities—a critical step for harmonizing data from diverse sources. Extracted treatment data were compared against expert-annotated ground truth. Model performance was assessed using accuracy metrics (Precision, Recall, F1-Score) and opera- tional feasibility metrics, including processing speed and structural compliance of the output. Results: A strong positive correlation was observed between model size and extraction accuracy. F1-score improved from 0.609 for the 1B model to 0.710 (3B), 0.807 (8B), and 0.828 (70B). While larger models demonstrated superior accuracy and compliance, they incurred higher computational costs. The modest performance difference between 8B and 70B suggests diminishing returns with increasing model size. Conclusions: LLMs represent a viable technology for automating oncology treatment extraction. The 8B-parameter model emerged as a highly effective option, balancing high accuracy and computational efficiency. Selecting an appropriate LLM for deployment in cancer registries involves a trade-off between desired accuracy and available operational resources. Harmonizing extracted entities with the oncology knowledge base facilitates standardized integration into common data models, enhancing data quality for real-world evidence analyses.

artificial intelligence↗

Post-fire soil respiration in late growing season (2023 and 2024), Kougarok Fire Complex, Seward Peninsula, Alaska

Field soil respiration data collected in 2023 and 2024 from burned and unburned tussock tundra sites in the Kougarok Fire Complex, near Nome, on the Seward Peninsula of Alaska. Specifically, we measured soil properties and late-growing season CO2 fluxes in patches of unique plant functional types (forbs, shrubs, and graminoids) across two years in tundra recovering from repeated wildfires over the decade. The goal was to identify the main drivers of soil respiration in Arctic tundra underlain by discontinuous permafrost that is recovering from two recent, repeated wildfires that differed in fire age and number of times burned, thereby resulting in different levels of vegetation and subsurface property changes (i.e., successional trajectories). There are five files in *.csv format with one data file and four data description files including data, dictionary, methods, terminology, and file-level metadata. The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic), is a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research. The NGEE Arctic Phase 3 project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska. Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

Santos, Fernanda [ORNL] (ORCID:0000000191555623)↗

NGEE Arctic Integrated Modeling (IM3): Improved snow-vegetation interaction

This data product represents the integration of new code capability for arctic tundra snow-vegetation-terrain interactions into the Energy Exascale Earth System Model (E3SM), through the E3SM Land Model (ELM) component. This code integration is the result of collaborative effort between the NGEE Arctic project and the E3SM project. The NGEE Arctic project developed a total of six Integrated Modeling (IM) modules informed by observations and experiments. New ELM capability represented by this data product (IM3) falls into three categories: 1) Downscaling from gridcell to topographic unit level when working through the existing coupler bypass code. 2) Four new parameters (taper, stocking, bendresist, and vegshape) have been added to ELM to allow for flexible definition of snow-vegetation interactions. 3) Vegshape and bendresist parameters are used to calculate the fraction of leaf area and/or stem area buried by snow for a given snow depth. This data record consists of a single document (pdf format) that describes the theoretical basis for the snow-vegetation-terrain interactions added to ELM, and describes the modifications made to the ELM code. The Methods section of this metadata record includes a link to the public E3SM code repository where the exact code modifications as integrated in E3SM can be accessed. The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic), was a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research. The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska. Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

Thornton, Peter E [ORNL] (ORCID:0000000247595158)↗

NGEE Arctic Integrated Modeling (IM2): Improved subgrid hillslope hydrologic connectivity

This data product represents the integration of new code capability for arctic tundra hillslope hydrologic processes into the Energy Exascale Earth System Model (E3SM), through the E3SM Land Model (ELM) component. This code integration is the result of collaborative effort between the NGEE Arctic project and the E3SM project. The current ELM represents water movement primarily through vertical processes, such as precipitation, canopy interception, evaporation, infiltration, and soil water movement. Lateral water movement—such as surface runoff, subsurface flow, and river transport—plays a significant role in the hydrological cycle, especially in regions with varied topography. While E3SM includes a runoff routing component representing water transport in the river network, the lateral transport of water at the subgrid scale within the land model has previously not been taken into account. With the recent development of topographic units within the ELM subgrid data structure, there is an opportunity to simulate hillslope hydrologic connectivity by introducing water transport along topographic gradients. We expect that more realistic representation of hillslope hydrologic processes will lead to improved predictions of both soil water content and river network flows. Lateral transport of water at and near the surface is represented as a sub-grid process in this new code development. Water is tracked as it moves from higher to lower elevations within a gridcell. This capability uses the nested hierarchical sub-grid scheme within ELM to connect water fluxes from sub-grid elements with higher elevation to those with lower elevation. This data record consists of a single document (pdf format) that describes the theoretical basis for the hillslope hydrology processes added to ELM, and describes the modifications made to the ELM code. The Methods section of this metadata record includes a link to the public E3SM code repository where the exact code modifications as integrated in E3SM can be accessed. The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic), was a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research. The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska. Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

Thornton, Peter E [ORNL] (ORCID:0000000247595158)↗

Intelligent Manufacturing Support: Specialized LLMs for Composite Material Processing and Equipment Operation

Engineering educational curriculum and standards cover many material and manufacturing options. However, engineers and designers are often unfamiliar with certain composite materials or manufacturing techniques. Large language models (LLMs) could potentially bridge the gap. Their capacity to store and retrieve data from large databases provides them with a breadth of knowledge across disciplines. However, their generalized knowledge base can lack targeted, industry-specific knowledge. To this end, we present two LLM-based applications based on the GPT-4 architecture: (1) The Composites Guide: a system that provides expert knowledge on composites material and connects users with research and industry professionals who can provide additional support and (2) The Equipment Assistant: a system that provides guidance for manufacturing tool operation and material characterization. By combining the knowledge of general AI models with industry-specific knowledge, both applications are intended to provide more meaningful information for engineers. In this paper, we discuss the development of the applications and evaluate it through a benchmark and two informal user studies. The benchmark analysis uses the Rouge and Bertscore metrics to evaluate our models’ performance against GPT-4o. The results show that GPT-4o and the proposed models perform similarly or better on the ROUGE and BERTScore metrics. The two user studies supplement this quantitative evaluation by asking experts to provide qualitative and open-ended feedback about our model’s performance on a set of domain-specific questions. The results of both studies highlight a potential for more detailed and specific responses with the Composites Guide and the Equipment Assistant.

Kapoor, Gunnika [Oak Ridge National Laboratory (OR↗

Quantum Computing for AI-based Design and Optimization of Electric Motors

Knowledge-based artificial intelligence and hierarchical fuzzy logic offer an interpretable framework for electricvehicle motor preliminary design, but their computational burden grows with linguistic granularity and coupled design-space size. This paper presents a reduced quantum reformulation of the hierarchical fuzzy inference of air-gap flux density, a representative level-one motor-design parameter. Starting from the published electric-vehicle motor-design framework, a three-term fuzzy prototype is constructed from the original inference structure. The reduced model is then reformulated as a modular quantum register-oracle system, in which each hierarchical subrelation is encoded as a block oracle and evaluated through superpositionbased candidate-label testing. The proposed modular quantum formulation reproduces the reduced classical prototype after block fusion. A resource analysis shows that the reduced modular system requires seven qubits per block and twenty-two qubits in a straightforward four-block implementation. Finally, a crossovercomplexity model is derived to identify the regime in which quantum candidate search may become favorable relative to hierarchical fuzzy inference. The results show that no quantum advantage should be claimed for the present one-output reduced benchmark, but that a plausible crossover emerges for larger joint candidate spaces and higher linguistic granularity. The work therefore establishes a technically consistent starting point for future quantum-assisted electric-vehicle motor-design optimization.

Kumar, Praveen [ORNL] (ORCID:0000000291877857)↗

KEBLM: Knowledge-Enhanced Biomedical Language Models

Pretrained language models (PLMs) have demonstrated strong performance on many natural language processing (NLP) tasks. Despite their great success, these PLMs are typically pretrained only on unstructured free texts without leveraging existing structured knowledge bases that are readily available for many domains, especially scientific domains. As a result, these PLMs may not achieve satisfactory performance on knowledge-intensive tasks such as biomedical NLP. Comprehending a complex biomedical document without domain-specific knowledge is challenging, even for humans. Inspired by this observation, we propose a general framework for incorporating various types of domain knowledge from multiple sources into biomedical PLMs. We encode domain knowledge using lightweight adapter modules, bottleneck feed-forward networks that are inserted into different locations of a backbone PLM. For each knowledge source of interest, we pretrain an adapter module to capture the knowledge in a self-supervised way. We design a wide range of self-supervised objectives to accommodate diverse types of knowledge, ranging from entity relations to description sentences. Once a set of pretrained adapters is available, we employ fusion layers to combine the knowledge encoded within these adapters for downstream tasks. Each fusion layer is a parameterized mixer of the available trained adapters that can identify and activate the most useful adapters for a given input. Our method diverges from prior work by including a knowledge consolidation phase, during which we teach the fusion layers to effectively combine knowledge from both the original PLM and newly-acquired external knowledge using a large collection of unannotated texts. After the consolidation phase, the complete knowledge-enhanced model can be fine-tuned for any downstream task of interest to achieve optimal performance. Extensive experiments on many biomedical NLP datasets show that our proposed framework consistently improves the performance of the underlying PLMs on various downstream tasks such as natural language inference, question answering, and entity linking. These results demonstrate the benefits of using multiple sources of external knowledge to enhance PLMs and the effectiveness of the framework for incorporating knowledge into PLMs. Finally, while primarily focused on the biomedical domain in this work, our framework is highly adaptable and can be easily applied to other domains, such as the bioenergy sector.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Machine learning guided optimal composition selection of niobium alloys for high temperature applications

Nickel- and cobalt-based superalloys are commonly used as turbine materials for high-temperature applications. However, their maximum operating temperature is limited to about 1100 °C. Therefore, to improve turbine efficiency, current research is focused on designing materials that can withstand higher temperatures. Niobium-based alloys can be considered as promising candidates because of their exceptional properties at elevated temperatures. The conventional approach to alloy design relies on phase diagrams and structure–property data of limited alloys and extrapolates this information into unexplored compositional space. In this work, we harness machine learning and provide an efficient design strategy for finding promising niobium-based alloy compositions with high yield and ultimate tensile strength. Unlike standard composition-based features, we use domain knowledge-based custom features and achieve higher prediction accuracy. We apply Bayesian optimization to screen out novel Nb-based quaternary and quinary alloy compositions and find these compositions have superior predicted strength over a range of temperatures. We develop a detailed design flow and include Python programming code, which could be helpful for accelerating alloy design in a limited alloy data regime.

Mohanty, Trupti (ORCID:0000000342701430)↗

Atomic layer deposition of nanofilms on porous polymer substrates: Strategies for success

Atomic layer deposition (ALD) is a versatile technique for engineering the surfaces of porous polymers, imbuing the flexible, high-surface-area substrates with inorganic and hybrid material properties. Previously reported enhancements include fouling resistance, electrical conductance, thermal stability, photocatalytic activity, hydrophilicity, and oleophilicity. However, there are many poorly understood phenomena that introduce challenges in applying ALD to porous polymers. In this paper, we address five common challenges and ways to overcome them: (1) entrapped precursor, (2) embrittlement, (3) film fracture, (4) deformation, and (5) pore collapse. These challenges are often interrelated and can exacerbate one another. To investigate these phenomena, we applied various ALD chemistries to porous polymers including polyethersulfone, polysulfone, polyvinylidene fluoride, and polycarbonate track-etched membranes. Reaction-diffusion modeling revealed why certain precursors and processing conditions result in embrittling subsurface material growth, entrapment of unreacted precursors, and nongrowth. We quantify the limits of ALD processing temperatures that are dictated by thermal expansion mismatch and can lead to fractured ALD films. The results herein allow us to make recommendations to avoid, mitigate, or overcome the difficulties encountered when performing ALD and plasma-enhanced ALD on porous polymers. We intend this article to serve as a “lessons learned” guide informed by previous experience to provide a better understanding of the difficulties and limitations of ALD on porous polymers and knowledge-based guidelines for successful depositions. This knowledge can accelerate future research and help experimentalists navigate and troubleshoot as they expose porous polymers to reactive precursor vapors.

36 MATERIALS SCIENCE↗

BETO 2021 Peer Review - FCIC Task 6: High Temperature Conversion

The impacts of feedstock variability on pyrolysis processes are significant but poorly defined. Current engineering designs are based on empirical guidelines, useful only over a narrow range of feedstock properties. The objectives of this project are to (1) Develop science-based knowledge of how feedstock attributes and operational parameters impact pyrolysis process reliability and product quality; and (2) Build an experimental and computational toolset that predicts these outcomes, enabling processes to optimize reliability and product quality. Biomass is a complex feedstock. Controlling for and testing the effects of individual attributes is very challenging. This project couples multiscale experimentation and modeling to accurately capture the fundamental physics and chemistry of biomass flow and conversion behavior in feeding and pyrolysis reactor operations. Our focus is on pine residue attributes – anatomical fraction (bark, needles, wood), particle morphology (size/shape distribution, density, porosity), and chemical composition (extractives, biopolymers, alkali metals) – that impact product quality for downstream catalytic upgrading. Because detailed pyrolysis product characterization is limited, cutting-edge analytical techniques are being developed to reveal impactful product attributes. The tools and knowledge developed here will enable integrated pyrolysis-based processes that are more robust, flexible, and market-responsive with respect to feedstock variability.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Quantifying Firebrand Production and Transport Using the Acoustic Analysis of In-Fire Cameras

Firebrand travel and ignition of spot fires is a major concern in the Wildland-Urban Interface and in wildfire operations overall. Firebrands allow for the efficient breaching across fuel-free barriers such as roads, rivers and constructed fuel breaks. Existing observation-based knowledge on medium-distance firebrand travel is often based on single tree experiments that do not replicate the intensity and convective updraft of a continuous crown fire. Recent advances in acoustic analysis, specifically pattern detection, has enabled the quantification of the rate at which firebrands are observed in the audio recordings of in-fire cameras housed within fire-proof steel boxes that have been deployed on experimental fires. The audio pattern being detected is the sound created by a flying firebrand hitting the steel box of the camera. This technique allows for the number of firebrands per second to be quantified and can be related to the fire's location at that same time interval (using a detailed rate of spread reconstruction) in order to determine the firebrand travel distance. A proof of concept is given for an experimental crown fire that shows the viability of this technique. When related to the fire's location, key areas of medium-distance spotting are observed that correspond to regions of peak fire intensity. Trends on the number of firebrands landing per square metre as the fire approaches are readily quantified using low-cost instrumentation.

58 GEOSCIENCES↗

Knowledge-guided graph machine learning for spatially distributed prediction of daily discharge and nitrogen export dynamics

Spatially distributed prediction of streamflow and nitrogen export dynamics is essential for precision management of agricultural watersheds. While temporal deep learning models such as Long Short-Term Memory (LSTM) have shown strong performance at basin scales, their ability to generalize spatially is limited by insufficient representation of spatial dependencies and flow paths, particularly under data-scarce conditions. To address this gap, we propose HydroGraphNet, a knowledge-guided graph machine learning framework that integrates process-based knowledge and explicit spatial learning into temporal modeling. This framework incorporates directed graph topology to encode watershed connectivity and upstream inflows, with mass balance constraints to improve physical consistency. To enhance generalization in sparsely monitored regions, HydroGraphNet is pretrained on synthetic data generated by the SWAT+ (Soil and Water Assessment Tool Plus) model. We evaluated HydroGraphNet in the Upper Sangamon River Basin (44 HUC-12 subwatersheds, 2001–2020) against two LSTM baselines: a lumped basin-level model and a distributed variant. When benchmarked on SWAT+ simulations in pretraining, HydroGraphNet improved test NSEs by 8.9% (discharge) and 13.7% (NO₃–N load) in temporal extrapolation, and by 27.1% and 34.7% in spatial extrapolation, relative to the Lumped LSTM baseline. After fine-tuning with USGS monitoring data, the model achieved mean test NSE (KGE) scores of 0.768 (0.861) for discharge and 0.626 (0.664) for NO₃–N load, substantially outperforming baselines. Attribution analysis further highlighted the importance of upstream inflow representation and graph-based spatial learning in capturing cross-subwatershed dependencies. The model also reproduced seasonal hydrological and biogeochemical patterns consistent with known processes, demonstrating its robustness and process fidelity for spatially distributed prediction. Altogether, HydroGraphNet advances the integration of physical knowledge and spatially explicit learning in hydrological modeling, offering a generalizable framework for distributed modeling to support spatially targeted water quality management in data-scarce watersheds.

54 ENVIRONMENTAL SCIENCES↗

Knowledge-guided learning with curated prior genetic biomarkers for robust model interpretation

Abstract Motivation Knowledge-guided learning offers effective and robust model training strategies in data-scarce settings by incorporating established domain knowledge, thereby enhancing generalization, robustness, and interpretability. By contrast, conventional deep learning approaches rely purely on data-driven learning, which can limit robust model interpretability, particularly in high-dimensional settings with limited size samples. In computational biology, knowledge-guided learning has primarily leveraged network- and structural-based knowledge, leading to biologically interpretable representations and enhanced predictive performance compared to conventional approaches. However, curated biomarkers, one of the most accessible forms of biological knowledge, remain largely unexplored within knowledge-guided paradigms. Results In this study, we propose a model-agnostic training paradigm, Biomarker-driven Explainable Prior-guided Learning (BioExPL), that can be applied to any neural networks that incorporates curated prior knowledge. BioExPL enforces neural networks to reflect curated biomarker priors in their latent representations through a novel knowledge-alignment loss. BioExPL consistently demonstrated significantly improved predictive performance and enhanced model interpretability with minimized computational overhead in simulation studies and intensive experiments on multiple cancer datasets. BioExPL not only integrates prior curated knowledge into the model but also accurately identifies unknown associated signals additionally. BioExPL is model-agnostic and domain-independent, enabling its integration into diverse neural network architectures. Availability and implementation The open-source is publicly available at: https://github.com/datax-lab/BioExPL.

Baek, Beomsu [Department of Computer Science, Univ↗

Advanced Algal Biofoundries for the Production of Polyurethane Precursors [Abstract]

The collaboration between PNNL, LBNL, and UCSD will develop novel algae bioproduction platforms for the generation of polymer precursors. To achieve this, we will design and construct genetic tools for enhanced production of target chemicals. Biosensors will be developed to detect target chemical production. We will expand our knowledge of metabolic and regulatory systems that control or inhibit biosynthesis of the target molecules, and use this knowledge base to optimize production titers. Finally, we will integrate the knowledge gained from each activity to produce a set of the most promising production strains used to demonstrate the economic viability of producing polymer precursors in algal bioproduction platforms at scale under either heterotrophic or photosynthetic growth.

60 APPLIED LIFE SCIENCES↗