Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “knowledge representation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Knowledge Beacons: Web services for data harvesting of distributed biomedical knowledge

The continually expanding distributed global compendium of biomedical knowledge is diffuse, heterogeneous and huge, posing a serious challenge for biomedical researchers in knowledge harvesting: accessing, compiling, integrating and interpreting data, information and knowledge. In order to accelerate research towards effective medical treatments and optimizing health, it is critical that efficient and automated tools for identifying key research concepts and their experimentally discovered interrelationships are developed. As an activity within the feasibility phase of a project called “Translator” (https://ncats.nih.gov/translator) funded by the National Center for Advancing Translational Sciences (NCATS) to develop a biomedical science knowledge management platform, we designed a Representational State Transfer (REST) web services Application Programming Interface (API) specification, which we call a Knowledge Beacon. Knowledge Beacons provide a standardized basic API for the discovery of concepts, their relationships and associated supporting evidence from distributed online repositories of biomedical knowledge. This specification also enforces the annotation of knowledge concepts and statements to the NCATS endorsed the Biolink Model data model and semantic encoding standards (https://biolink.github.io/biolink-model/). Implementation of this API on top of diverse knowledge sources potentially enables their uniform integration behind client software which will facilitate research access and integration of biomedical knowledge.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Chemical reaction enhanced graph learning for molecule representation

Abstract Motivation Molecular representation learning (MRL) models molecules with low-dimensional vectors to support biological and chemical applications. Current methods primarily rely on intrinsic molecular information to learn molecular representations, but they often overlook effectively integrating domain knowledge into MRL. Results In this article, we develop a reaction-enhanced graph learning (RXGL) framework for MRL, utilizing chemical reactions as domain knowledge. RXGL introduces dual graph learning modules to model molecule representation. One module employs graph convolutions on molecular graphs to capture molecule structures. The other module constructs a reaction-aware graph from chemical reactions and designs a novel graph attention network on this graph to integrate reaction-level relations into molecular modeling. To refine molecule representations, we design a reaction-based relation learning task, which considers the relations between the reactant and product sides in reactions. In addition, we introduce a cross-view contrastive task to strengthen the cooperative associations between molecular and reaction-aware graph learning. Experiment results show that our RXGL achieves strong performance in various downstream tasks, including product prediction, reaction classification, and molecular property prediction. Availability and implementation The code is publicly available at https://github.com/coder-ACAC/RLM.

Biochemistry & Molecular Biology↗

Leveraging Large Language Models for Understanding Fundamental Principles of Catalysis

Heterogeneous catalysis presents a distinct challenge for artificial intelligence (AI). Data sets are often small and inconsistently reported, catalyst representations are not standardized, and extracting fundamental knowledge requires integrating performance data, spectroscopic characterizations, and mechanistic models across multiple scales. Language offers a unifying representation across these modalities, making catalysis well suited for leveraging large language models (LLMs). By standardizing how catalytic data is represented, LLMs make dispersed experimental results more accessible to downstream statistical modeling. In this perspective, we focus our discussion around three opportunities where LLMs can significantly contribute to catalysis: (1) text to properties; (2) text to structure; and (3) text to mechanistic models. The discussion is followed by a perspective section on LLM-readiness of data, aligning LLM outputs with scientific correctness, and bridging lab-scale discovery to industrial deployment. Across each area, the most productive applications couple dispersed chemical knowledge with physics-grounded validation to produce verifiable hypotheses and actionable representations.

Catalysts↗

ARCH: Large-scale knowledge graph via aggregated narrative codified health records analysis

Objective: Electronic health record (EHR) systems contain a wealth of clinical data stored as both codified data and free-text narrative notes (NLP). The complexity of EHR presents challenges in feature representation, information extraction, and uncertainty quantification. Here, to address these challenges, we proposed an efficient Aggregated naRrative Codified Health (ARCH) records analysis to generate a large-scale knowledge graph (KG) for a comprehensive set of EHR codified and narrative features. Methods: Using data from 12.5 million Veterans Affairs patients, ARCH first derives embedding vectors and generates similarities along with associated p-values to measure the strength of relatedness between clinical features with statistical certainty quantification. Next, ARCH performs a sparse embedding regression to remove indirect linkage between features to build a sparse KG. Finally, ARCH was validated on various clinical tasks, including detecting known relationships between entity pairs, predicting drug side effects, disease phenotyping, as well as sub-typing Alzheimer’s disease patients. Results: ARCH produces high-quality clinical embeddings and KG for over 60,000 codified and narrative EHR concepts. The KG and embeddings are visualized in the R-shiny powered web-API.3 ARCH achieved high accuracy in detecting EHR concept relationships, with AUCs of 0.926 (codified) and 0.861 (NLP) for similar EHR concepts, and 0.810 (codified) and 0.843 (NLP) for related pairs. It detected drug side effects with a 0.723 AUC, which improved to 0.826 after fine-tuning. Using both codified and NLP features, the detection power increased significantly. Compared to other methods, ARCH has superior accuracy and enhances weakly supervised phenotyping algorithms’ performance. Notably, it successfully categorized Alzheimer’s patients into two subgroups with varying mortality rates. Conclusion: The proposed ARCH algorithm generates large-scale high-quality semantic representations and knowledge graph for both codified and NLP EHR features, useful for a wide range of predictive modeling tasks.

Electronic health records↗

Microbially mediated climate feedbacks from wetland ecosystems

Wetlands are crucial nodes in the carbon cycle, emitting approximately 20% of global CH4 while also sequestering 20%–30% of all soil carbon. Both greenhouse gas fluxes and carbon storage are driven by microbial communities in wetland soils. However, these key players are often overlooked or overly simplified in current global climate models. Here, we first integrate microbial metabolisms with biological, chemical, and physical processes occurring at scales from individual microbial cells to ecosystems. This conceptual scale-bridging framework guides the development of feedback loops describing how wetland-specific climate impacts (i.e., sea level rise in estuarine wetlands, droughts and floods in inland wetlands) will affect future climate trajectories. Furthermore, these feedback loops highlight knowledge gaps that need to be addressed to develop predictive models of future climates capturing microbial contributions. We propose a roadmap connecting environmental scientific disciplines to address these knowledge gaps and improve the representation of microbial processes in climate models. Together, this paves the way to understand how microbially mediated climate feedbacks from wetlands will impact future climate change.

54 ENVIRONMENTAL SCIENCES↗

Understanding the Physics Representation of Deep Learning Models in Environmental Applications

Deep learning (DL) models have been popular in earth and environmental modeling and analysis, which exhibit huge potential in capturing and reconstructing the non-linearity of relevant environmental processes. They are extensively used as analytical tools or emulators for multiple domains (atmosphere, land surface, ocean, and biogeochemistry). Despite their success, their internal working mechanism remains largely unknown. Such a lack of knowledge hinders the identification of physically consistent models that are fully adaptive to non-stationary climate, as well as the development of physics-informed machine learning such as physics-informed neural network (PINN). To establish preliminary knowledge and framework of such physics representation evaluation, this project focuses on an improved understanding of DL models in the environmental applications. DL models are increasingly applied to environmental modeling and prediction. However, they have been evaluated mostly from a performance perspective, and there is a gap in understanding how they represent the known physics internally. Such knowledge is especially critical when applying DL models under climate change conditions, where new inputs are likely outside the ranges of the training datasets. In this project, we reveal how the known physical processes are represented within DL models from both statistical and mechanistic perspectives. Leveraging the traditional model evaluations that focus more on the accuracies of predictions, we establish a framework that examines both the accuracy and physics representation of DL models. This analysis framework can identify DL models that make the correct predictions based on correct physics, thus enhancing the existing explainable artificial intelligence (explainable-AI) portfolio. It lays a foundation for developing novel metrics to evaluate the emerging DL models in environmental applications. This knowledge also informs the development of physics-informed DL models by revealing the direct connections between the known physical processes and specific model components or structures.

54 ENVIRONMENTAL SCIENCES↗

Algebraic model to study the internal structure of pseudoscalar mesons with heavy-light quark content

The internal structure of all lowest-lying pseudoscalar mesons with heavy-light quark content is studied in detail using an algebraic model that has been applied recently, and successfully, to the same physical observables of pseudoscalar and vector mesons with hidden-flavor quark content, from light to heavy quark sectors. The algebraic model consists on constructing simple and evidence-based of the meson’s Bethe-Salpeter amplitude (BSA) and quark’s propagator in such a way that the Bethe-Salpeter wave function (BSWF) can then be readily computed algebraically. Its subsequent projection onto the light front yields the light front wave function (LFWF) whose form allows us a simple access to the valence-quark parton distribution amplitude (PDA) by integrating over the transverse momentum squared. We exploit our current knowledge of the PDAs of lowest-lying pseudoscalar heavy-light mesons to compute their generalized parton distributions (GPDs) through the overlap representation of LFWFs. From these three dimensional knowledge, different limits/projections lead us to deduce the related parton distribution functions (PDFs), electromagnetic form factors (EFFs), and impact parameter space GPDs (IPS-GPDs). When possible, we make explicit comparisons with available experimental results and earlier theoretical predictions. Published by the American Physical Society 2024

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

DOME: Directional medical embedding vectors from Electronic Health Records

Motivation: The increasing availability of Electronic Health Record (EHR) systems has created enormous potential for translational research. Recent developments in representation learning techniques have led to effective large-scale representations of EHR concepts along with knowledge graphs that empower downstream EHR studies. However, most existing methods require training with patient-level data, limiting their abilities to expand the training with multi-institutional EHR data. On the other hand, scalable approaches that only require summary-level data do not incorporate temporal dependencies between concepts. Methods: We introduce a DirectiOnal Medical Embedding (DOME) algorithm to encode temporally directional relationships between medical concepts, using summary-level EHR data. Specifically, DOME first aggregates patient-level EHR data into an asymmetric co-occurrence matrix. Then it computes two Positive Pointwise Mutual Information (PPMI) matrices to correspondingly encode the pairwise prior and posterior dependencies between medical concepts. Following that, a joint matrix factorization is performed on the two PPMI matrices, which results in three vectors for each concept: a semantic embedding and two directional context embeddings. They collectively provide a comprehensive depiction of the temporal relationship between EHR concepts. Results: We highlight the advantages and translational potential of DOME through three sets of validation studies. First, DOME consistently improves existing direction-agnostic embedding vectors for disease risk prediction in several diseases, for example achieving a relative gain of 5.5% in the area under the receiver operating characteristic (AUROC) for lung cancer. Second, DOME excels in directional drug-disease relationship inference by successfully differentiating between drug side effects and indications, correspondingly achieving relative AUROC gain over the state-of-the-art methods by 10.8% and 6.6%. Finally, DOME effectively constructs directional knowledge graphs, which distinguish disease risk factors from comorbidities, thereby revealing disease progression trajectories. The source codes are provided at https://github.com/celehs/Directional-EHRembedding.

60 APPLIED LIFE SCIENCES↗

Robustness of topological persistence in knowledge distillation for wearable sensor data

Topological data analysis (TDA) has shown great success in various applications involving wearable sensor data. However, there are difficulties in leveraging topological features in machine learning and wearable sensors because of the large time consumption and computational resources required to extract the features. To address this problem, knowledge distillation (KD) is utilized to generate a small model and accommodate topological features with persistence image (PI) representations from the raw time series data. Deploying topological knowledge in KD enables the student to achieve better performance compared to the one trained solely on raw time series data. However, it is not yet known if there are coherent characteristics for topological features in PI, which can aid in improving the performance during KD. In this paper, we investigate the suitability and challenges of utilizing topological features in KD for wearable sensor data, thereby contributing to the advancement of the field. Our study explores the impact of transferred topological features by comparing the Teacher-to-Student framework with Multiple Teachers-to-Student where teachers utilize both time series data and persistence images obtained by TDA as inputs. Additionally, we conduct a rigorous examination of topological knowledge effects by testing under various corruptions, knowledge types, and learning strategies in the context of human activity recognition tasks. Our analysis of topological features in KD presents the optimal strategy for incorporating these features. This study includes datasets of varying scales, window lengths, and activity classes, providing a comprehensive evaluation. Our results demonstrate that leveraging topological features in KD to enhance performance across databases.

97 MATHEMATICS AND COMPUTING↗

Conflation of Geospatial POI Data and Ground-level Imagery

The code solves the problem of conflating POI (Points of Interest) geospatial data and geo-tagged ground-level imagery data. The main challenges with fusion or conflation of these two types of data has been that these two data entities are represented not only in several ways in current state-of-the-art, but also there is mismatch of representation format of these two data types. This source code/software brings POI datapoints and ground-level imagery datapoints into same representation format, and then generates valuable Knowledge Graph combining POI and images data (so that it can be used for multitude of applications).

De, Debraj↗

A change language for ontologies and knowledge graphs

Ontologies and knowledge graphs (KGs) are general-purpose computable representations of some domain, such as human anatomy, and are frequently a crucial part of modern information systems. Most of these structures change over time, incorporating new knowledge or information that was previously missing. Managing these changes is a challenge, both in terms of communicating changes to users and providing mechanisms to make it easier for multiple stakeholders to contribute. To fill that need, we have created KGCL, the Knowledge Graph Change Language (https://github.com/INCATools/kgcl), a standard data model for describing changes to KGs and ontologies at a high level, and an accompanying human-readable Controlled Natural Language (CNL). This language serves two purposes: a curator can use it to request desired changes, and it can also be used to describe changes that have already happened, corresponding to the concepts of “apply patch” and “diff” commonly used for managing changes in text documents and computer programs. Another key feature of KGCL is that descriptions are at a high enough level to be useful and understood by a variety of stakeholders—e.g. ontology edits can be specified by commands like “add synonym ‘arm’ to ‘forelimb’” or “move ‘Parkinson disease’ under ‘neurodegenerative disease’.” We have also built a suite of tools for managing ontology changes. These include an automated agent that integrates with and monitors GitHub ontology repositories and applies any requested changes and a new component in the BioPortal ontology resource that allows users to make change requests directly from within the BioPortal user interface. Overall, the KGCL data model, its CNL, and associated tooling allow for easier management and processing of changes associated with the development of ontologies and KGs.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Convoluted filtering for process cycle modeling

Principles of materials science and engineering, physics, mathematics, and information science are used to extract knowledge and insights from the process-structure–property-performance relationships hidden in materials data. The process-structure modeling can be accelerated without loss of interpretability, with artificial intelligence tools that mimic the salient features of the process and process-structure relations. In this work, a novel convoluted model-filtering technique was exploited to build and successfully train the Convoluted Filter (CoFi) artifacts for Fe-based alloy heat treatment cycles. The artifacts were pre-trained to filter out deep models that change the surrogate microstructure state after the heat treatment at ambient conditions. Direct representation of the thermal cycle features within knowledge Graph facilitated development of meaningful data models for microstructure evolution, which reduce overfitting to limited datasets.

36 MATERIALS SCIENCE↗

Data-Driven Refinement of Electronic Energies from Two-Electron Reduced-Density-Matrix Theory

The exponential computational cost of describing strongly correlated electrons can be mitigated by adopting a reduced-density-matrix (RDM)-based description of the electronic structure. While variational two-electron RDM (v2RDM) methods can enable large-scale calculations on such systems, the quality of the solution is limited by the fact that only a subset of known necessary N-representability constraints can be applied to the 2RDM in practical calculations. Here, we demonstrate that violations of partial three-particle (T1 and T2) N-representability conditions, which can be evaluated with knowledge of only the 2RDM, can serve as physics-based features in a machine-learning (ML) protocol for improving energies from v2RDM calculations that consider only two-particle (PQG) conditions. Proof-of-principle calculations demonstrate that the model yields substantially improved energies, relative to reference values from configuration-interaction-based calculations.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Characterization of carbonaceous aerosols during TRACER-CAT

Absorbing aerosols (AA) have an important impact on the global radiation budget and cloud properties. The composition and properties of AA can vary substantially throughout the atmosphere, depending on the particle source and the influence of chemical aging. Uncertainties associated with the radiative effects of AA remain substantial. A key contributor to this uncertainty is understanding the extent to which coatings in general, and water uptake especially, alters absorption by AA particles and how this depends on particle composition. We deployed new and existing experimental tools during the Tracking Aerosol Convection Interactions Experiment (TRACER) campaign in Houston, TX as part of the Carbonaceous Aerosols Thrust (CAT) to provide detailed characterization of aerosol optical, chemical, and physical properties. Our TRACER-CAT measurements complemented and expanded on the planned TRACER instrumentation, allowing for more detailed characterization of aerosol properties of relevance to cloud development (a core focus of TRACER), such as the composition of particles that can act as cloud condensation nuclei, than would otherwise be available. Our measurements have allowed for assessment of the relationship(s) between AA optical properties (with a focus on absorption) and the chemical and physical characteristics (including the mixing state of black carbon (BC) containing particles). These field observations occurred in collaboration with Los Alamos National Laboratory in summer 2022 during the TRACER intensive operating period. The instrumentation we co-deployed provided for measurement of (i) multi-wavelength dry aerosol absorption, scattering, and extinction, (ii) the size-dependent composition and abundance of sub-micron aerosol, differentiating between those particles that do and do not contain BC, (iii) BC-specific concentrations and size distributions, (iv) particle size, and (v) the first field measurements at an ARM site of the influence of RH on multi-wavelength absorption by ambient AA. We have leveraged the natural variability of the atmosphere and of aerosol sources in the Houston region to (i) specifically disentangle contributions to light absorption from BC, absorbing organic carbon (brown carbon), and coatings on BC, (ii) characterize the mixing state of BC and assess the factors that give rise to compositional differences between BC-containing and BC-free aerosol, and (iii) establish how water uptake influences absorption and how any such effect depends on particle composition and BC mixing state. Overall, our study contributed to the mission of the Atmospheric System Research program in multiple ways. Through the deployment of complementary, advanced instrumentation for characterization of a wide range of aerosol properties our work helped to maximize the scientific impact of the TRACER campaign. Our work also allowed for development of new insights into the relationship(s) between aerosol composition, hygroscopicity, and the mixing state of BC with aerosol optical properties. Through this, our work has provided knowledge that can improve understanding and model representation of aerosol processes as they affect the Earth’s radiation budget.

54 ENVIRONMENTAL SCIENCES↗

Transactional Knowledge Graph Generation To Model Adversarial Activities

A Knowledge Graph (KG) is a formal and structured representation of facts, relationships, and semantic descriptions of a set of entities. Traditionally, KGs are used to describe metadata about entities and to provide additional context to target application results. Many real-world domains also involve temporal interactions between entities in addition to the metadata data. Modeling these attributed transactions is a critical requirement when using KGs in complex real-world applications. Modeling adversarial activities is one such application that develops methodology and tools to produce realistic large-scale background activity graphs that include embedded Weapons of Mass Destruction (WMD) activity patterns. We present a novel platform for constructing a transactional knowledge graph from a diverse set of sources. We present the core components and architecture of the framework, and a use case for generating a background knowledge graph and WMD activity template to evaluate network alignment and subgraph matching algorithms.

Purohit, Sumit↗

COVID19 Disease Map, a computational knowledge repository of virus–host interaction mechanisms

We need to effectively combine the knowledge from surging literature with complex datasets to propose mechanistic models of SARS-CoV-2 infection, improving data interpretation and predicting key targets of intervention. Here, we describe a large-scale community effort to build an open access, interoperable and computable repository of COVID-19 molecular mechanisms. The COVID-19 Disease Map (C19DMap) is a graphical, interactive representation of disease-relevant molecular mechanisms linking many knowledge sources. Notably, it is a computational resource for graph-based analyses and disease modelling. To this end, we established a framework of tools, platforms and guidelines necessary for a multifaceted community of biocurators, domain experts, bioinformaticians and computational biologists. The diagrams of the C19DMap, curated from the literature, are integrated with relevant interaction and text mining databases. We demonstrate the application of network analysis and modelling approaches by concrete examples to highlight new testable hypotheses. This framework helps to find signatures of SARS-CoV-2 predisposition, treatment response or prioritisation of drug candidates. Such an approach may help deal with new waves of COVID-19 or similar pandemics in the long-term perspective.

59 BASIC BIOLOGICAL SCIENCES↗

Modeling dust mineralogical composition: sensitivity to soil mineralogy atlases and their expected climate impacts

Soil dust aerosols are a key component of the climate system, as they interact with short- and long-wave radiation, alter cloud formation processes, affect atmospheric chemistry and play a role in biogeochemical cycles by providing nutrient inputs such as iron and phosphorus. The influence of dust on these processes depends on its physicochemical properties, which, far from being homogeneous, are shaped by its regionally varying mineral composition. The relative amount of minerals in dust depends on the source region and shows a large geographical variability. However, many state-of-the-art Earth system models (ESMs), upon which climate analyses and projections rely, still consider dust mineralogy to be invariant. The explicit representation of minerals in ESMs is more hindered by our limited knowledge of the global soil composition along with the resulting size-resolved airborne mineralogy than by computational constraints. In this work we introduce an explicit mineralogy representation within the state-of-the-art Multiscale Online Nonhydrostatic AtmospheRe CHemistry (MONARCH) model. We review and compare two existing soil mineralogy datasets, which remain a source of uncertainty for dust mineralogy modeling and provide an evaluation of multiannual simulations against available mineralogy observations. Soil mineralogy datasets are based on measurements performed after wet sieving, which breaks the aggregates found in the parent soil. Our model predicts the emitted particle size distribution (PSD) in terms of its constituent minerals based on brittle fragmentation theory (BFT), which reconstructs the emitted mineral aggregates destroyed by wet sieving. Our simulations broadly reproduce the most abundant mineral fractions independently of the soil composition data used. Feldspars and calcite are highly sensitive to the soil mineralogy map, mainly due to the different assumptions made in each soil dataset to extrapolate a handful of soil measurements to arid and semi-arid regions worldwide. For the least abundant or more difficult-to-determine minerals, such as iron oxides, uncertainties in soil mineralogy yield differences in annual mean aerosol mass fractions of up to ~ 100 %. Although BFT restores coarse aggregates including phyllosilicates that usually break during soil analysis, we still identify an overestimation of coarse quartz mass fractions (above 2 µm in diameter). In a dedicated experiment, we estimate the fraction of dust with undetermined composition as given by a soil map, which makes up ~ 10 % of the emitted dust mass at the global scale and can be regionally larger. Changes in the underlying soil mineralogy impact our estimates of climate-relevant variables, particularly affecting the regional variability of the single-scattering albedo at solar wavelengths or the total iron deposited over oceans. All in all, this assessment represents a baseline for future model experiments including new mineralogical maps constrained by high-quality spaceborne hyperspectral measurements, such as those arising from the NASA Earth Surface Mineral Dust Source Investigation (EMIT) mission.

54 ENVIRONMENTAL SCIENCES↗

Qualitative trend analysis based on a mixed-integer representation

Shape constrained spline fitting is a useful method to impose prior knowledge onto flexible semi-parametric models during parameter estimation. Most typically, the function shape is imposed through order restrictions on the regression coefficients. The intended shape is considered known or selected based on heuristic rules. In this study, we present a method to estimate the optimal set of order restrictions to segment a univariate data series into episodes with distinct shapes. This is also known as the qualitative trend analysis (QTA) problem. The obtained solution uses a trade-off between lack-of-fit and model complexity. Further, our practical implementation takes inspiration from the generalized order restricted information criterion (GORIC) for inequality-constrained model selection. From this, one learns (a) that QTA can be formulated as a mixed-integer quadratic program (MIQP) and (b) that the newly proposed mixed order restricted information criterion (MORIC) enables optimal segmentation. This is illustrated through didactic case studies.

42 ENGINEERING↗