Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “mathematics of computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

HIResist: a database of HIV-1 resistance to broadly neutralizing antibodies

Changing the course of the human immunodeficiency virus type I (HIV-1) pandemic is a high public health priority with approximately 39 million people currently living with HIV-1 (PLWH) and about 1.5 million new infections annually worldwide. Broadly neutralizing antibodies (bnAbs) typically target highly conserved sites on the HIV-1 envelope glycoproteins (Envs), which mediate viral entry, and block the infection of diverse HIV-1 strains. But different mechanisms of HIV-1 resistance to bnAbs prevent robust application of bnAbs for therapeutic and preventive interventions. Here we report the development of a new database that provides data and computational tools to aid the discovery of resistant features and may assist in analysis of HIV-1 resistance to bnAbs. Bioinformatic tools allow identification of specific patterns in Env sequences of resistant strains and development of strategies to elucidate the mechanisms of HIV-1 escape; comparison of resistant and sensitive HIV-1 strains for each bnAb; identification of resistance and sensitivity signatures associated with specific bnAbs or groups of bnAbs; and visualization of antibody pairs on cross-sensitivity plots. The database has been designed with a particular focus on user-friendly and interactive interface. Our database is a valuable resource for the scientific community and provides opportunities to investigate patterns of HIV-1 resistance and to develop new approaches aimed to overcome HIV-1 resistance to bnAbs.

60 APPLIED LIFE SCIENCES↗

Reference-free structural variant detection in microbiomes via long-read co-assembly graphs

Motivation: The study of bacterial genome dynamics is vital for understanding the mechanisms underlying microbial adaptation, growth, and their impact on host phenotype. Structural variants (SVs), genomic alterations of 50 base pairs or more, play a pivotal role in driving evolutionary processes and maintaining genomic heterogeneity within bacterial populations. While SV detection in isolate genomes is relatively straightforward, metagenomes present broader challenges due to the absence of clear reference genomes and the presence of mixed strains. In response, our proposed method rhea, forgoes reference genomes and metagenome-assembled genomes (MAGs) by encompassing all metagenomic samples in a series (time or other metric) into a single co-assembly graph. The log fold change in graph coverage between successive samples is then calculated to call SVs that are thriving or declining. Results: We show rhea to outperform existing methods for SV and horizontal gene transfer (HGT) detection in two simulated mock metagenomes, particularly as the simulated reads diverge from reference genomes and an increase in strain diversity is incorporated. We additionally demonstrate use cases for rhea on series metagenomic data of environmental and fermented food microbiomes to detect specific sequence alterations between successive time and temperature samples, suggesting host advantage. Our approach leverages previous work in assembly graph structural and coverage patterns to provide versatility in studying SVs across diverse and poorly characterized microbial communities for more comprehensive insights into microbial gene flux.

59 BASIC BIOLOGICAL SCIENCES↗

Chemical reaction enhanced graph learning for molecule representation

Abstract Motivation Molecular representation learning (MRL) models molecules with low-dimensional vectors to support biological and chemical applications. Current methods primarily rely on intrinsic molecular information to learn molecular representations, but they often overlook effectively integrating domain knowledge into MRL. Results In this article, we develop a reaction-enhanced graph learning (RXGL) framework for MRL, utilizing chemical reactions as domain knowledge. RXGL introduces dual graph learning modules to model molecule representation. One module employs graph convolutions on molecular graphs to capture molecule structures. The other module constructs a reaction-aware graph from chemical reactions and designs a novel graph attention network on this graph to integrate reaction-level relations into molecular modeling. To refine molecule representations, we design a reaction-based relation learning task, which considers the relations between the reactant and product sides in reactions. In addition, we introduce a cross-view contrastive task to strengthen the cooperative associations between molecular and reaction-aware graph learning. Experiment results show that our RXGL achieves strong performance in various downstream tasks, including product prediction, reaction classification, and molecular property prediction. Availability and implementation The code is publicly available at https://github.com/coder-ACAC/RLM.

Biochemistry & Molecular Biology↗

CryoTEN: efficiently enhancing cryo-EM density maps using transformers

Abstract Motivation Cryogenic electron microscopy (cryo-EM) is a core experimental technique used to determine the structure of macromolecules such as proteins. However, the effectiveness of cryo-EM is often hindered by the noise and missing density values in cryo-EM density maps caused by experimental conditions such as low contrast and conformational heterogeneity. Although various global and local map-sharpening techniques are widely employed to improve cryo-EM density maps, it is still challenging to efficiently improve their quality for building better protein structures from them. Results In this study, we introduce CryoTEN—a 3D UNETR++ style transformer to improve cryo-EM maps effectively. CryoTEN is trained using a diverse set of 1295 cryo-EM maps as inputs and their corresponding simulated maps generated from known protein structures as targets. An independent test set containing 150 maps is used to evaluate CryoTEN, and the results demonstrate that it can robustly enhance the quality of cryo-EM density maps. In addition, automatic de novo protein structure modeling shows that protein structures built from the density maps processed by CryoTEN have substantially better quality than those built from the original maps. Compared to the existing state-of-the-art deep learning methods for enhancing cryo-EM density maps, CryoTEN ranks second in improving the quality of density maps, while running >10 times faster and requiring much less GPU memory than them. Availability and implementation The source code and data are freely available at https://github.com/jianlin-cheng/cryoten.

Biochemistry & Molecular Biology↗

miss-SNF: a multimodal patient similarity network integration approach to handle completely missing data sources

Abstract Motivation Precision medicine leverages patient-specific multimodal data to improve prevention, diagnosis, prognosis, and treatment of diseases. Advancing precision medicine requires the non-trivial integration of complex, heterogeneous, and potentially high-dimensional data sources, such as multi-omics and clinical data. In the literature, several approaches have been proposed to manage missing data, but are usually limited to the recovery of subsets of features for a subset of patients. A largely overlooked problem is the integration of multiple sources of data when one or more of them are completely missing for a subset of patients, a relatively common condition in clinical practice. Results We propose miss-Similarity Network Fusion (miss-SNF), a novel general-purpose data integration approach designed to manage completely missing data in the context of patient similarity networks. miss-SNF integrates incomplete unimodal patient similarity networks by leveraging a non-linear message-passing strategy borrowed from the SNF algorithm. miss-SNF is able to recover missing patient similarities and is “task agnostic”, in the sense that can integrate partial data for both unsupervised and supervised prediction tasks. Experimental analyses on nine cancer datasets from The Cancer Genome Atlas (TCGA) demonstrate that miss-SNF achieves state-of-the-art results in recovering similarities and in identifying patients subgroups enriched in clinically relevant variables and having differential survival. Moreover, amputation experiments show that miss-SNF supervised prediction of cancer clinical outcomes and Alzheimer’s disease diagnosis with completely missing data achieves results comparable to those obtained when all the data are available. Availability and implementation miss-SNF code, implemented in R, is available at https://github.com/AnacletoLAB/missSNF.

Biochemistry & Molecular Biology↗

NGPINT V3: a containerized orchestration Python software for discovery of next-generation protein–protein interactions

Abstract Summary Batch yeast two-hybrid (Y2H) assays, leveraged with next-generation sequencing, have afforded successful innovations for the analysis of protein–protein interactions. NGPINT is a Conda-based software designed to process the millions of raw sequencing reads resulting from Y2H–next-generation interaction screens. Over time, increasing compatibility and dependency issues have prevented clean NGPINT installation and operation. A system-wide update was essential to continue effective use with its companion software, Y2H-SCORES. We present NGPINT V3, a containerized implementation built with both Singularity and Docker, allowing accessibility across virtually any operating system and computing environment. Availability and implementation This update includes streamlined dependencies and container images hosted on Sylabs (https://cloud.sylabs.io/library/schuyler/ngpint/ngpint) and Dockerhub (https://hub.docker.com/r/schuylerds/ngpint), facilitating easier adoption and integration into high-throughput and cloud-computing workflows. Full instructions and software can be also found in the GitHub repository https://github.com/Wiselab2/NGPINT_V3 and Zenodo https://doi.org/10.5281/zenodo.15256036.

Biochemistry & Molecular Biology↗

CSGL: chemical synthesis graph learning for molecule representation

Abstract Motivation Molecule representation learning (MRL) translates molecules into a real vector space, serving as input to downstream tasks in biology, chemistry, and computer science. This article introduces a chemical synthesis graph learning (CSGL) framework, which enhances MRL by considering both the atomic structures of molecules and their roles in chemical reactions through a hierarchical graph representation. Specifically, molecules are first modeled based on their molecular graphs, which capture atomic-level structural information. They are then further refined using a chemical synthesis graph, where nodes represent reactant and product molecule sets, and edges encode chemical transformations between reactants and products (e.g. changes in molecular structures). CSGL optimizes molecular embeddings of reactant and product nodes in a fashion that ensures the embeddings conform to a chemical balance constraint. Results Experimental results show that our method CSGL achieves strong performance on a variety of tasks, including product prediction, reaction classification, and molecular property prediction. Availability and implementation https://github.com/li-2023/CSGL.

Biochemistry & Molecular Biology↗

SRBench++: Principled Benchmarking of Symbolic Regression With Domain-Expert Interpretation

Symbolic regression searches for analytic expressions that accurately describe studied phenomena. The main promise of this approach is that it may return an interpretable model that can be insightful to users, while maintaining high accuracy. The current standard for benchmarking these algorithms is SRBench, which evaluates methods on hundreds of datasets that are a mix of real-world and simulated processes spanning multiple domains. At present, the ability of SRBench to evaluate interpretability is limited to measuring the size of expressions on real-world data, and the exactness of model forms on synthetic data. In practice, model size is only one of many factors used by subject experts to determine how interpretable a model truly is. Furthermore, SRBench does not characterize algorithm performance on specific, challenging sub-tasks of regression such as feature selection and evasion of local minima. In this work, we propose and evaluate an approach to benchmarking SR algorithms that addresses these limitations of SRBench by 1) incorporating expert evaluations of interpretability on a domain-specific task, and 2) evaluating algorithms over distinct properties of data science tasks. We evaluate 12 modern symbolic regression algorithms on these benchmarks and present an in-depth analysis of the results, discuss current challenges of symbolic regression algorithms and highlight possible improvements for the benchmark itself.

97 MATHEMATICS AND COMPUTING↗

Integrated Framework of Multisource Data Fusion for Outage Location in Looped Distribution Systems

Accurate outage location is essential for expediting post-outage power restoration, minimizing outage duration, and enhancing the resilience of distribution networks. With the advent of advanced metering infrastructure, data-driven outage location methods have significantly advanced beyond traditional approaches that rely on manual inspections. However, existing methods still face critical challenges, like reliance on single-source data, limited ability to handle partially observable systems or difficulties with loop networks. To the best of our knowledge, no single approach has comprehensively addressed all of these challenges at once. To this end, this paper proposes a comprehensive multisource data fusion framework for outage locations via probabilistic graph networks. The framework consists of three key phases. First, a novel method for reconstituting distribution networks with loops is developed, transforming looped networks into multiple radial subnetworks that retain all outage causalities of the original network. Second, Bayesian network (BN) models are established for each subnetwork, integrating multiple data sources and network structures. Finally, a joint Gibbs sampling mechanism, featuring forward and backward information flow, is designed to merge data from separate BN models and maximize the utilization of limited evidence, ensuring accurate outage location identification. In conclusion, the framework was validated on two modified public test systems, and comparative studies confirmed its effectiveness.

24 POWER TRANSMISSION AND DISTRIBUTION↗

The Eight Levers of Coercive Conflict: No More, No Less

Few concepts are more important to our nation than the principle of deterrence. It is at the very core of our national security strategy. Despite its lasting import to our nation, and the rest of the world, we still lack a complete and explicit formulation of the calculus of deterrence. This has, at times, resulted in failed policies costing the Nation lives and wealth. With the ultimate objective of making the Nation more secure, we establish a complete and explicit formulation of deterrence calculus, which consists of eight levers. As it turns out this formulation also applies to the calculus of compellence, and is thus a unifying model of coercive conflict.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

A mathematical framework for thermodynamic computing with applications to chemical reaction networks

The widespread adoption of energy-intensive computing applications has led to a growing need for energy-efficient computing approaches. Thermodynamic computing offers a promising approach for low-energy computation by leveraging the intrinsic computational capabilities of physical, chemical, or biological systems. However, the mathematical foundations of thermodynamic computing require further development to fully realize the potential energy efficiencies, as well as to assess factors like noise and operational speed. In this paper, we establish a mathematical framework for utilizing thermodynamic processes to perform fundamental operations, including addition, subtraction, multiplication, and division. We highlight the use of chemical reactions as potential computational units and explore synthetic chemical and biochemical systems as practical implementations. Additionally, we demonstrate how these principles can be applied to solving complex mathematical problems, such as ordinary differential equations (ODEs) and suggest the necessary components to implement the thermodynamic computing framework using chemical reactions based in a microfluidic device. This work enhances our understanding of thermodynamic processes for natural computing as a basis for scalable, energy-efficient computation in paradigm disruptive next-generation systems.

Cannon, William R. [Pacific Northwest National Lab↗

Likelihood Maximization and Moment Matching in Low SNR Gaussian Mixture Models

We derive an asymptotic expansion for the log-likelihood of Gaussian mixture models (GMMs) with equal covariance matrices in the low signal-to-noise regime. The expansion reveals an intimate connection between two types of algorithms for parameter estimation: the method of moments and likelihood optimizing algorithms such as Expectation-Maximization (EM). We show that likelihood optimization in the low SNR regime reduces to a sequence of least squares optimization problems that match the moments of the estimate to the ground truth moments one by one. This connection is a stepping stone towards the analysis of EM and maximum likelihood estimation in a wide range of models. A motivating application for the study of low SNR mixture models is cryo-electron microscopy data, which can be modeled as a GMM with algebraic constraints imposed on the mixture centers. We discuss the application of our expansion to algebraically constrained GMMs, among other example models of interest. © 2022 The Authors. Communications on Pure and Applied Mathematics published by Wiley Periodicals LLC.

97 MATHEMATICS AND COMPUTING↗

Estimating posterior quantity of interest expectations in a multilevel scalable framework

Scalable approaches for uncertainty quantification are necessary for characterizing prediction confidence in large-scale subsurface flow simulations with uncertain permeability. To this end we explore a multilevel Monte Carlo approach for estimating posterior moments of a particular quantity of interest, where we employ an element-agglomerated algebraic multigrid (AMG) technique to generate the hierarchy of coarse spaces with guaranteed approximation properties for both the generation of spatially correlated random fields and the forward simulation of Darcy's law to model subsurface flow. In both these components (sampling and forward solves), we exploit solvers that rely on state-of-the-art scalable AMG. To illustrate the applicability of this approach, numerical tests are performed on two 3D examples-a unit cube and an egg-shaped domain with an irregular boundary-where the scalability of each simulation as well as the scalability of the overall algorithm are demonstrated.

97 MATHEMATICS AND COMPUTING↗

Convergence analysis of single rate and multirate fixed stress split iterative coupling schemes in heterogeneous poroelastic media

Recently, the accurate modeling of flow–structure interactions has gained more attention and importance for both petroleum and environmental engineering applications. Of particular interest is the coupling between subsurface flow and reservoir geomechanics. Different single rate and multirate iterative and explicit coupling schemes have been proposed and analyzed in the past. In addition, Banach fixed point contraction results were obtained for iterative coupling schemes, and conditionally stable results were obtained for explicit coupling schemes. In this work, we will consider the mathematical analysis of the single rate and multirate fixed stress split iterative coupling schemes for spatially heterogeneous poroelastic media. We will re–establish the contractivity for both schemes in the localized case, and we will show that heterogeneities come at the expense of imposing more restricted conditions on the number of fine flow time steps that can be taken within one coarse mechanics time step in the multirate case. Our mathematical analysis is supplemented by numerical simulations validating our derived upper bounds. Finally, to the best of our knowledge, this is the first rigorous mathematical analysis of the multirate fixed–stress split iterative coupling scheme in heterogeneous poroelastic media.

97 MATHEMATICS AND COMPUTING↗

Nearly Periodic Maps and Geometric Integration of Noncanonical Hamiltonian Systems

Abstract M. Kruskal showed that each continuous-time nearly periodic dynamical system admits a formal U (1)-symmetry, generated by the so-called roto-rate. When the nearly periodic system is also Hamiltonian, Noether’s theorem implies the existence of a corresponding adiabatic invariant. We develop a discrete-time analog of Kruskal’s theory. Nearly periodic maps are defined as parameter-dependent diffeomorphisms that limit to rotations along a U (1)-action. When the limiting rotation is non-resonant, these maps admit formal U (1)-symmetries to all orders in perturbation theory. For Hamiltonian nearly periodic maps on exact presymplectic manifolds, we prove that the formal U (1)-symmetry gives rise to a discrete-time adiabatic invariant using a discrete-time extension of Noether’s theorem. When the unperturbed U (1)-orbits are contractible, we also find a discrete-time adiabatic invariant for mappings that are merely presymplectic, rather than Hamiltonian. As an application of the theory, we use it to develop a novel technique for geometric integration of non-canonical Hamiltonian systems on exact symplectic manifolds.

97 MATHEMATICS AND COMPUTING↗