Weapons Knowledge Management Correcting a Knowledge Misconception
Explore the source record for details and available documents.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
There has been a recent push to increase access to electric vehicle (EV) charging infrastructure. The National Electric Vehicle Infrastructure (NEVI) program, part of the Bipartisan Infrastructure Law (BIL) has made significant funding available for major charging infrastructure projects along state thruways, and many state and local incentives exist for EV owners to install chargers in their homes. However, deployment of these chargers has not kept up with demand, primarily due to issues in project planning, permitting processes, and unforeseen delays. This paper serves as a review of the current understanding of these and other non-hardware costs in EV charging infrastructure projects (collectively known as “soft costs”). We found that soft costs in EV charging infrastructure projects are not well understood. Specifically, there is little agreement on how soft costs should be categorized and tracked, and less agreement still on best practices for controlling these costs and lowering barriers to infrastructure deployment. A broader review of EV charging infrastructure cost analyses shows that these costs can have significant impacts on project outcomes. EV charging infrastructure projects may be able to examine the success of the solar industry in lowering soft costs, and a similar effort may lower project costs significantly. Further work on standardizing and collecting data on EV charging infrastructure costs is required to begin addressing and controlling these costs.
Not Available
Abstract Motivation Knowledge-guided learning offers effective and robust model training strategies in data-scarce settings by incorporating established domain knowledge, thereby enhancing generalization, robustness, and interpretability. By contrast, conventional deep learning approaches rely purely on data-driven learning, which can limit robust model interpretability, particularly in high-dimensional settings with limited size samples. In computational biology, knowledge-guided learning has primarily leveraged network- and structural-based knowledge, leading to biologically interpretable representations and enhanced predictive performance compared to conventional approaches. However, curated biomarkers, one of the most accessible forms of biological knowledge, remain largely unexplored within knowledge-guided paradigms. Results In this study, we propose a model-agnostic training paradigm, Biomarker-driven Explainable Prior-guided Learning (BioExPL), that can be applied to any neural networks that incorporates curated prior knowledge. BioExPL enforces neural networks to reflect curated biomarker priors in their latent representations through a novel knowledge-alignment loss. BioExPL consistently demonstrated significantly improved predictive performance and enhanced model interpretability with minimized computational overhead in simulation studies and intensive experiments on multiple cancer datasets. BioExPL not only integrates prior curated knowledge into the model but also accurately identifies unknown associated signals additionally. BioExPL is model-agnostic and domain-independent, enabling its integration into diverse neural network architectures. Availability and implementation The open-source is publicly available at: https://github.com/datax-lab/BioExPL.
KOGUT — Knowledge Oriented Graph Unified Transformer KOGUT implements the Relational Graph Transformer (RelGT) architecture for knowledge graph link prediction in biological domains, with a primary focus on microbial growth media prediction. While the original RelGT (arXiv:2505.10960) targets relational tables, time series, and multi-table databases, KOGUT adapts this architecture for heterogeneous biological knowledge graphs, providing first-in-class AI predictive models for microbial cultivation. Key Adaptations Beyond Original RelGT: - Knowledge Graph Focus: Applied to biological KGs with semantic node types (taxa, chemicals, media, phenotypes, environments) versus generic relational database tables, trained on the KG-Microbe knowledge graph (1.3M entities, 2.9M edges, 24 relation types). - Multimodal Node Encoding: Integrates node labels, categories, descriptions, and synonyms from KG metadata through learned embedding layers—adapting relational column features to graph node attributes with textual semantics. - Extended K-Hop Subgraph Strategy: Optimized neighborhood sampling (3-hop default, configurable up to 200 nodes) tuned for sparse biological networks, building on the original local-global attention framework with biological relation preservation. - Biolink Predicate Preservation: Type-specific transformations for 24 biological edge semantics (occurs_in, consumes, produces, has_phenotype, subclass_of) beyond standard relational foreign keys, enabling multi-relation link prediction. - Inductive Learning Support: Enables zero-shot predictions for novel taxa through feature-based embeddings (temperature, oxygen requirements, gram stain, cell shape), extending the original transductive relational benchmark scope to uncultured microorganisms. CheapSOTA Performance Optimizations (This Distribution): - VQ-EMA Centroid Attention: Vector quantization with exponential moving average for improved global context modeling (+5-10% MRR improvement). - HDF5 Precomputed Data Loading: One-time preprocessing of k-hop subgraphs to eliminate redundant graph traversals (2-5× training speedup). - Distributed Data Parallel Training: Multi-GPU support for scaling to larger knowledge graphs (tested on 4× NVIDIA A100 GPUs at NERSC Perlmutter). - Mixed Precision Training: Automatic mixed precision (AMP) for memory efficiency and faster training. Advantages Over Standard Knowledge Graph Embedding Models: Combines RelGT's proven multi-element tokenization (features, type, hop, structure) with graph-native biological representations, enabling interpretable link prediction across heterogeneous entities that standard embedding models (TransE, RotatE, ComplEx) and table-based transformers cannot directly model. Achieves near-perfect performance on microbial growth media prediction (MRR: 0.9966, Precision@1: 0.9932, Hit@10: 1.0000) while maintaining explainability through attention-based reasoning over biological pathways. Training Data: - KG-Microbe merged knowledge graph: 1,379,337 nodes, 2,960,472 edges - 24 biological relation types including taxonomic hierarchies, metabolic interactions, phenotype associations, and environmental relationships - Primary prediction task: Growth media suitability for microbial taxa (biolink:occurs_in, 50K edges) - Multi-relation capability: Predicts links for any of the 24 relation types, including chemical consumption/production, phenotype associations, and taxonomic classification Citation: Original RelGT Architecture: Dwivedi et al., "Relational Graph Transformer", arXiv:2505.10960, 2025 KOGUT Implementation: Knowledge Oriented Graph Unified Transformer for Microbial Growth Media Prediction Developed at Lawrence Berkeley National Laboratory (LBNL) Trained on NERSC Perlmutter supercomputer
This report explores using "generator knowledge" to determine whether a solar photovoltaic (PV) module must be managed as hazardous waste prior to recycling or landfilling. Generator knowledge is a legal term and existing regulatory pathway for making a hazardous waste determination that has been used by other industries but is a relatively unknown option to the PV industry. In the United States, a hazardous waste determination often acts as a pre-requisite to recycle or landfill a PV module. The results of the hazardous waste determination dictate whether the PV module must be managed as hazardous waste or nonhazardous solid waste. Managing a PV module as hazardous waste requires compliance with stringent U.S. federal and state hazardous waste law. In addition to increased management costs, which can be ten times higher, legal liability for PV modules regulated as hazardous waste is also heightened with both civil and criminal penalties for noncompliance which includes making an inaccurate or faulty hazardous waste determination. The most common reason a PV module would be regulated as hazardous is if it contains a regulated metal in an amount that equals or exceeds the toxicity characteristic limits. To determine whether a PV module exhibits a hazardous characteristic, the regulated person/entity must "apply knowledge...in light of the materials and processes used." In the absence of adequate knowledge to determine whether the PV module is hazardous, it must be tested using Test Method 1311 Toxicity Characteristic Leaching Procedure (TCLP) or an equivalent EPA-approved method. Although TCLP is the predominant method used to today to make a hazardous waste determination for PV modules, evidence from this study concludes it is not practical to TCLP test every PV module even in a single utility-scale installation, and a scalable solution is needed. This study finds that knowledge-based hazardous waste determinations may allow a regulated person/entity to make a hazardous waste determination for more than one PV module at one time - making this regulatory pathway a potential scalable solution. Through legal analysis and interviews with 44 experts, the authors explore what it means to make an accurate knowledge-based hazardous determination for PV modules considering sources and forms of information as well as potential limitations. The work aims to provide a foundation for building consensus on whether knowledge-based hazardous determinations are a viable, scalable industry approach for solar.
In order to guarantee that a system meets adequate levels of reliability and availability, system performances are continuously monitored and analyzed thanks to the technological advancements driving the Industry 4.0 revolution. An Industry 4.0 approach is typically based on advanced statistical, big data mining, machine learning, and internet-of-things methods designed to detect anomalies in the behavior of system, detect the most likely failure modes, and provide indications to system engineers on when maintenance activities should be performed before system performance are deemed unacceptable (which can be generated by diagnostic and prognostic methods). However, these analyses, which are designed to automatize and increase the efficacy of the system maintenance program, require large amount of data which can come in various forms: numeric, textual, images, sounds etc. Such data constitutes the historic knowledge benchmark to track system performances and support system engineer decisions. Here we claim that data is not sufficient to support this kind of analyses when applied to systems characterized by complex architectures and behaviors. Robust system engineer decisions require the ability to understand the system operational context that lies behind the observed data elements. In this respect, system models are in fact necessary to “put data in context” and capture relationships between data elements. Industry 4.0 methods require in fact contextual knowledge as a basis upon which hypotheses can be generated and assumptions tested. In our view, for complex systems, model-based system engineering (MBSE) models can afford this contextual knowledge, as they are typically used to describe systems architecture and dynamic behaviors. System knowledge is here intended as the blending of collected data and system architecture which takes the form of a “knowledge graph”. A knowledge graph is a database which consists of a large set of nodes (in our case an entity can be either a data or an MBSE element) which are linked to each other. The types of nodes and links follow a pre-defined topology, sometimes also refers as an ontology, that is designed to fit the actual decisions that needs to be performed. We show here how a knowledge graph can be defined to support system engineer maintenance decisions and how the same graph can be built based on system MBSE models and pre-processed data from numeric (through anomaly detections and diagnostic methods) and textual elements (through technical language processing TLP).
Nuclear safeguards were first announced in 1945 by the U.S., Canadian and British governments as a means to exchange scientific information about peaceful uses of atomic energy and prevent the use of nuclear material for weapons. International nuclear safeguards, now under the provision of the International Atomic Energy Agency (IAEA), serves to hold accountable the 140 member countries (States) that have entered into treaties and agreements against the spread of nuclear weapons. Most notably, the IAEA acts as a nuclear materials inspectorate under the global Nuclear Non-Proliferation Treaty, brought about in 1968, that specifies commitments member States make to the world?s non-proliferation regime. The Office of International Nuclear Safeguards (OINS) is part of the National Nuclear Security Administration, a semi-autonomous agency within the U.S. Department of Energy. Within OINS, the Human Capital Development program recognizes the need to build workforce capacity and safeguards expertise. This includes supporting the education and training of younger generations working in international nuclear safeguards. One of the program?s main concerns is core competencies being lost to retirement or attrition and building knowledge retention pipelines from senior to new professionals. One identified core competency is knowledge gained by individuals who have completed IAEA assignments such as missions to member States to conduct safeguards inspections. Such individuals possess unique knowledge, and efforts are being put forth to effectively capture this knowledge within the national laboratories complex. To this end, we report findings from a one-year mentor-mentee knowledge retention program in which an international safeguards subject matter expert imparted knowledge and skills to a willing professional mentee. The knowledge and skills stem from the mentor?s multi-year assignment at the IAEA in Vienna, Austria.
Abstract Challenges associated with global change stressors on ecosystems have prompted calls to improve actionable science, including through boundary‐spanning activities, which aim to build connections and communication between researchers and natural resource practitioners. By synthesizing and translating research and practitioner knowledge, boundary‐spanning activities could support proactive, research‐informed conservation practice, but the success of these efforts is rarely evaluated. Using repeat survey data from the Northeast Regional Invasive Species and Climate Change (NE RISCC) Management Network, a boundary‐spanning organization, we evaluate whether participating in NE RISCC affected practitioners' knowledge, actions and priorities related to invasive species management under a changing climate. Our survey results suggest that practitioners who participate in NE RISCC have greater knowledge about invasive species and climate change and are incorporating climate change in more ways into their invasive species management. We also found NE RISCC membership affected the perceived usefulness of informational resources, with NE RISCC members more frequently identifying research syntheses and targeted workshops (both are common products used by NE RISCC to translate science into practice and share manager knowledge) as useful compared to non‐members. Practitioners who participate in NE RISCC also identified somewhat different research priorities, with non‐members and short‐term members more frequently identifying range‐shifting neonative species and their impacts on native communities as higher priorities compared to long‐term NE RISCC members. NE RISCC research activities and outreach materials have consistently framed range‐shifting neonative species as comparatively low risk, suggesting that this information has influenced practitioner's perception of risk. Practical implication : Although real‐world impacts of applied ecology are notoriously difficult to quantify, this analysis illustrates that if research results are actively translated, they can affect the knowledge and actions of natural resource practitioners. These impacts illustrate the potential for boundary‐spanning efforts to address other global change challenges to conservation.
Topological data analysis (TDA) has shown great success in various applications involving wearable sensor data. However, there are difficulties in leveraging topological features in machine learning and wearable sensors because of the large time consumption and computational resources required to extract the features. To address this problem, knowledge distillation (KD) is utilized to generate a small model and accommodate topological features with persistence image (PI) representations from the raw time series data. Deploying topological knowledge in KD enables the student to achieve better performance compared to the one trained solely on raw time series data. However, it is not yet known if there are coherent characteristics for topological features in PI, which can aid in improving the performance during KD. In this paper, we investigate the suitability and challenges of utilizing topological features in KD for wearable sensor data, thereby contributing to the advancement of the field. Our study explores the impact of transferred topological features by comparing the Teacher-to-Student framework with Multiple Teachers-to-Student where teachers utilize both time series data and persistence images obtained by TDA as inputs. Additionally, we conduct a rigorous examination of topological knowledge effects by testing under various corruptions, knowledge types, and learning strategies in the context of human activity recognition tasks. Our analysis of topological features in KD presents the optimal strategy for incorporating these features. This study includes datasets of varying scales, window lengths, and activity classes, providing a comprehensive evaluation. Our results demonstrate that leveraging topological features in KD to enhance performance across databases.
The "RNA world" represents a novel frontier for the study of fundamental biological processes and human diseases and is paving the way for the development of new drugs tailored to each patient's biomolecular characteristics. Although scientific data about coding and non-coding RNA molecules are constantly produced and available from public repositories, they are scattered across different databases and a centralized, uniform, and semantically consistent representation of the "RNA world" is still lacking. We propose RNA-KG, a knowledge graph (KG) encompassing biological knowledge about RNAs gathered from more than 60 public databases, integrating functional relationships with genes, proteins, and chemicals and ontologically grounded biomedical concepts. To develop RNA-KG, we first identified, pre-processed, and characterized each data source; next, we built a meta-graph that provides an ontological description of the KG by representing all the bio-molecular entities and medical concepts of interest in this domain, as well as the types of interactions connecting them. Finally, we leveraged an instance-based semantically abstracted knowledge model to specify the ontological alignment according to which RNA-KG was generated. RNA-KG can be downloaded in different formats and also queried by a SPARQL endpoint. A thorough topological analysis of the resulting heterogeneous graph provides further insights into the characteristics of the "RNA world". RNA-KG can be both directly explored and visualized, and/or analyzed by applying computational methods to infer bio-medical knowledge from its heterogeneous nodes and edges. The resource can be easily updated with new experimental data, and specific views of the overall KG can be extracted according to the bio-medical problem to be studied.
Ontologies and knowledge graphs (KGs) are general-purpose computable representations of some domain, such as human anatomy, and are frequently a crucial part of modern information systems. Most of these structures change over time, incorporating new knowledge or information that was previously missing. Managing these changes is a challenge, both in terms of communicating changes to users and providing mechanisms to make it easier for multiple stakeholders to contribute. To fill that need, we have created KGCL, the Knowledge Graph Change Language (https://github.com/INCATools/kgcl), a standard data model for describing changes to KGs and ontologies at a high level, and an accompanying human-readable Controlled Natural Language (CNL). This language serves two purposes: a curator can use it to request desired changes, and it can also be used to describe changes that have already happened, corresponding to the concepts of “apply patch” and “diff” commonly used for managing changes in text documents and computer programs. Another key feature of KGCL is that descriptions are at a high enough level to be useful and understood by a variety of stakeholders—e.g. ontology edits can be specified by commands like “add synonym ‘arm’ to ‘forelimb’” or “move ‘Parkinson disease’ under ‘neurodegenerative disease’.” We have also built a suite of tools for managing ontology changes. These include an automated agent that integrates with and monitors GitHub ontology repositories and applies any requested changes and a new component in the BioPortal ontology resource that allows users to make change requests directly from within the BioPortal user interface. Overall, the KGCL data model, its CNL, and associated tooling allow for easier management and processing of changes associated with the development of ontologies and KGs.
Motivation: Predicting microbial gene fitness across environmental conditions remains a central challenge for predictive phenomics and autonomous experimentation. Fitness assays generate large volumes of genotype–phenotype measurements difficult to integrate with experimental metadata and biological function in a form that supports mechanistic reasoning. Knowledge graphs offer a semantic framework for unifying modalities and enabling context-aware inference. Results: We build GIMME (Graph Inference for Microbial Metabolism Exploration), a semantically grounded knowledge graph that unifies gene fitness measurements spanning 10 Pseudomonas species with experimental metadata and biological context. Media are decomposed into chemical components and experiments carry structured links to natural-language descriptions. The resulting graph supports two inference modes: (1) symbolic graph traversal to surface candidate gene–environment and gene–chemical associations, and (2) learned inference using heterogeneous graph neural networks that propagate information across neighborhoods. We formulate link regression over (gene, media, experiment) triplets, combining learned gene embeddings with pretrained LLM sourced text embeddings of node descriptions to predict gene fitness. We then augment a baseline MLP with an auxiliary message-passing encoder (GraphSAGE/GAT) that propagates information over gene–protein–function and media–chemical subgraphs, and fuse the two pathways with a gated residual connection. This approach produces strong agreement with held-out fitness measurements (GraphSAGE Pearson r 0.74) while also highlighting inference challenges in extreme-fitness regimes. We aggregate GAT edge-attention weights by relation type and layer to estimate which biological and environmental relations most influence fitness predictions. Conclusion: This work explores using knowledge graphs as “context graphs” for microbial phenotype prediction. They provide a rich substrate which enables explainable retrieval of supporting evidence, and provides a natural bridge to autonomous workflows that prioritize the next experiment.
Spatially distributed prediction of streamflow and nitrogen export dynamics is essential for precision management of agricultural watersheds. While temporal deep learning models such as Long Short-Term Memory (LSTM) have shown strong performance at basin scales, their ability to generalize spatially is limited by insufficient representation of spatial dependencies and flow paths, particularly under data-scarce conditions. To address this gap, we propose HydroGraphNet, a knowledge-guided graph machine learning framework that integrates process-based knowledge and explicit spatial learning into temporal modeling. This framework incorporates directed graph topology to encode watershed connectivity and upstream inflows, with mass balance constraints to improve physical consistency. To enhance generalization in sparsely monitored regions, HydroGraphNet is pretrained on synthetic data generated by the SWAT+ (Soil and Water Assessment Tool Plus) model. We evaluated HydroGraphNet in the Upper Sangamon River Basin (44 HUC-12 subwatersheds, 2001–2020) against two LSTM baselines: a lumped basin-level model and a distributed variant. When benchmarked on SWAT+ simulations in pretraining, HydroGraphNet improved test NSEs by 8.9% (discharge) and 13.7% (NO₃–N load) in temporal extrapolation, and by 27.1% and 34.7% in spatial extrapolation, relative to the Lumped LSTM baseline. After fine-tuning with USGS monitoring data, the model achieved mean test NSE (KGE) scores of 0.768 (0.861) for discharge and 0.626 (0.664) for NO₃–N load, substantially outperforming baselines. Attribution analysis further highlighted the importance of upstream inflow representation and graph-based spatial learning in capturing cross-subwatershed dependencies. The model also reproduced seasonal hydrological and biogeochemical patterns consistent with known processes, demonstrating its robustness and process fidelity for spatially distributed prediction. Altogether, HydroGraphNet advances the integration of physical knowledge and spatially explicit learning in hydrological modeling, offering a generalizable framework for distributed modeling to support spatially targeted water quality management in data-scarce watersheds.
Nuclear theory and experiments, alongside astrophysical observations, constrain the equation of state (EOS) of supranuclear-dense matter. Conversely, knowledge of the EOS allows an improved interpretation of nuclear or astrophysical data. In this article, we use several established constraints on the EOS and the new NICER measurement of PSR J0437-4715 to comment on the nature of the primary companion in GW230529 and the companion of PSR J0514-4002E. We find that, with a probability of ≳84% and ≳68%, respectively, both objects are black holes. These likelihoods increase to above 95% when one uses GW170817’s remnant as an upper limit on the TOV mass. We also demonstrate that the current knowledge of the EOS substantially disfavors high masses and radii for PSR J0030+0451, inferred recently when combining NICER with XMM-Newton background data and using particular hot-spot models. Lastly, we also use our obtained EOS knowledge to comment on measurements of the nuclear symmetry energy, finding that the large value predicted by the PREX-II measurement displays some mild tension with other constraints on the EOS.
Identifying gene targets for enhancing metabolite production in metabolic engineering is challenging due to the vast research literature and the approximation in genome-scale metabolic model (GEM) simulations. Here, to address this, we propose the Gene-Metabolite Association Prediction task, which automates gene discovery for given metabolite-gene pairs, accompanied by a benchmark dataset of 2474 metabolites and 1947 genes for Saccharomyces cerevisiae (SC) and Issatchenkia orientalis (IO). This task is complicated by incomplete metabolic graphs and metabolic heterogeneity. We introduce an Interactive Knowledge Transfer mechanism based on Metabolism Graphs (IKT4Meta) to enhance prediction accuracy by integrating cross-metabolism knowledge. Using Pretrained Language Models (PLMs) to generate inter-graph links mitigates heterogeneity issues, while intra-graph links are propagated via these anchors. Gene-metabolite predictions are then performed on the enriched graphs integrating multiple microorganisms’ knowledge. Experiments show that IKT4Meta outperforms baselines by up to 12.3% in link prediction.
Per- and polyfluoroalkyl substances (PFAS) have been a rising concern for the past two decades, with the United States Department of Defense and Environmental Protection Agency investing millions of dollars in research into remediation and clean-up technologies. Due to the environmental persistence, toxicity, biological uptake, and ongoing changes in both federal and state regulatory space, understanding the fate and transport of PFAS compounds has been of growing concern to the US Department of Energy (DOE). The DOE’s Hanford Site is investigating historical use of PFAS and will be doing site characterization for PFAS. Thus, PFAS have not yet been identified as a contaminant concern in regulatory documents. Based on historical records that mention the discharge of aqueous film-forming foam containing PFAS and having on-site fire stations (a risk factor for PFAS contamination), it seems likely that environmental releases of PFAS may have occurred. Pump and treat (P&T) remediation is the selected remedy for multiple groundwater contaminant plumes at Hanford. These P&T systems use ion exchange (IX) as a component of aboveground treatment, with the specific resins depending on the target contaminants. There is potential that these IX resins may be able to remove PFAS from groundwater, but investigation is needed to understand affinity/selectivity and removal capacity given the groundwater composition and the operating conditions. This report provides background on PFAS uses and chemistry, then provides a review of IX resin applications for PFAS, identifying knowledge gaps. Recommendations are provided regarding research needed to address knowledge gaps and acquire information needed to propose IX as a future PFAS remediation technology at the Hanford Site, as well as other U.S. Department of Energy sites. Generally, PFAS compounds are fluorinated substances that contain at least one fully fluorinated methyl or methylene carbon – with a few noted exceptions, any chemical with at least a perfluorinated methyl group (–CF3) or a perfluorinated methylene group (–CF2–) is a PFAS. These chemical compounds are characterized as non-biodegradable, non-reactive, non-photolytic, and hydrolysis resistant. This makes them highly recalcitrant within the environment, however polyfluoroalkyl materials are less recalcitrant as the carbon chains contain C–H bonds which are more easily broken than carbon – fluorine (C–F) bonds. The backbone carbon structures are commonly punctuated with a head group, the most well-known of them are perfluorooctanesulfonic acid and perfluorooctanoic acid, which possess a sulfonate and a carboxylate group, respectively. IX resins are marketed for the removal of PFAS from water systems and industrial water, however, the mechanism of removal is not as well understood as for anion or cation removal. A better understanding of the mechanism of removal would enable the development of IX resins that have improved specificity for PFAS removal. Four knowledge gaps were identified: 1) the effect of dissolved ions on the IX resin PFAS removal effectiveness, 2) the effect of additional primary contaminants of concern (PCOCs) or secondary contaminants of concern (SCOCs) on the effectiveness of PFAS via IX resin, 3) the mechanisms of PFAS removal from water, and 4) practical solutions to IX resin regeneration and waste disposal.
Plant bioengineering has generated tens of thousands of genotype-to-phenotype relationships, but this knowledge remains fragmented across narrative literature and difficult to use computationally. Inconsistent descriptions of DNA constructs, host species, and traits, including variable species names, omitted regulatory elements, and inconsistent gene symbols, impede data reuse, comparative analysis, and design-build-test-learn cycles. Here, we present the Plant Bioengineering Atlas, a literature-mined, ontology-grounded knowledge base assembled using an artificial intelligence (AI)-aided extraction pipeline. A large language model parsed open-access primary research articles to generate structured, provenance-anchored records of engineered genes, modification types, promoter-gene-terminator constructs, host species, target traits, and reported phenotypes, with every record traceable to its source. The current release contains 14,358 curated records encompassing 6,998 distinct genes across 436 plant species from 6,452 papers published between 2000 and 2026. Corpus analysis reveals that experiments are concentrated in a small group of model and crop species, disease and pathogen resistance is the most frequently engineered trait class, and constitutive regulatory parts (particularly the CaMV 35S promoter and NOS terminator) remain pervasive. Two in five records omit one or both flanking regulatory elements (i.e., promoter and terminator), while only 23.4% describe cassettes in which both elements resolve to named part classes, exposing a systematic reproducibility gap. We organize these data into a knowledge graph linking genes, constructs, species, and traits; provide access through an interactive web portal; and propose an AI-compatible documentation standard for AI-ready reporting. The Plant Bioengineering Atlas provides a foundation for data-driven hypothesis generation and AI-aided plant biodesign.