Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “consensus algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

A Hierarchical Framework for CO2 Storage Capacity in Deep Saline Aquifer Formations

Carbon dioxide (CO 2 ) storage in deep saline aquifers is a vital option for CO 2 mitigation at a large scale. Determining storage capacity is one of the crucial steps toward large-scale deployment of CO 2 storage. Results of capacity assessments tend toward a consensus that sufficient resources are available in saline aquifers in many parts of the world. However, current CO 2 capacity assessments involve significant inconsistencies and uncertainties caused by various technical assumptions, storage mechanisms considered, algorithms, and data types and resolutions. Furthermore, other constraint factors (such as techno-economic features, site suitability, risk, regulation, social-economic situation, and policies) significantly affect the storage capacity assessment results. Consequently, a consensus capacity classification system and assessment method should be capable of classifying the capacity type or even more related uncertainties. We present a hierarchical framework of CO 2 capacity to define the capacity types based on the various factors, algorithms, and datasets. Finally, a review of onshore CO 2 aquifer storage capacity assessments in China is presented as examples to illustrate the feasibility of the proposed hierarchical framework.

58 GEOSCIENCES↗

Adversarial Binaries: AI-guided Instrumentation Methods for Malware Detection Evasion

Adversarial binaries are executable files that have been altered without loss of function by an AI agent in order to deceive malware detection systems. Progress in this emergent vein of research has been constrained by the complex and rigid structure of executable files. Although prior work has demonstrated that these binaries deceive a variety of malware classification models which rely on disparate feature sets, a consensus as to the best approach has not been reached, either in terms of the optimization algorithms or the instrumentation methods. Furthermore, although inconsistencies in the data sets, target classifiers, and functionality verification methods make head-to-head comparisons difficult, here we extract lessons learned and make recommendations for future research.

malware obfuscation↗

From Points to Planes: A Workflow for Converting Three‐Dimensional Point Cloud Data Into Discrete Fracture Network Flow and Transport Models

We present the Point cLoud Algorithm for NEtwork Extraction of Discrete Fracture Networks (PLANE-DFN), a point cloud–based algorithm for automatic fracture network extraction designed to support discrete fracture network (DFN) modeling workflows. PLANE-DFN segments three-dimensional fracture planes from raw point cloud data using RANdom SAmple Consensus coupled with statistical outlier removal and density-based clustering to isolate individual fracture features. Each candidate plane is constrained against site-specific structural constraints based on strike and dip. After segmentation, each fracture is converted into a 2-D convex polygon suitable for meshing and simulation. The PLANE-DFN algorithm is validated by comparing geometric and flow and transport data against data from dfnWorks simulations with ensembles of plane-fit networks. We find that the flow and transport in plane-fit networks are comparable to dfnWorks-generated networks when realistic network geometry is maintained. The PLANE-DFN algorithm provides an automated and streamlined workflow to transform point clouds of data into DFN network geometry.

54 ENVIRONMENTAL SCIENCES↗

Machine learning models for rat multigeneration reproductive toxicity prediction

Reproductive toxicity is one of the prominent endpoints in the risk assessment of environmental and industrial chemicals. Due to the complexity of the reproductive system, traditional reproductive toxicity testing in animals, especially guideline multigeneration reproductive toxicity studies, take a long time and are expensive. Therefore, machine learning, as a promising alternative approach, should be considered when evaluating the reproductive toxicity of chemicals. We curated rat multigeneration reproductive toxicity testing data of 275 chemicals from ToxRefDB (Toxicity Reference Database) and developed predictive models using seven machine learning algorithms (decision tree, decision forest, random forest, k-nearest neighbors, support vector machine, linear discriminant analysis, and logistic regression). A consensus model was built based on the seven individual models. An external validation set was curated from the COSMOS database and the literature. The performances of individual and consensus models were evaluated using 500 iterations of 5-fold cross-validations and the external validation data set. The balanced accuracy of the models ranged from 58% to 65% in the 5-fold cross-validations and 45%–61% in the external validations. Prediction confidence analysis was conducted to provide additional information for more appropriate applications of the developed models. The impact of our findings is in increasing confidence in machine learning models. We demonstrate the importance of using consensus models for harnessing the benefits of multiple machine learning models (i.e., using redundant systems to check validity of outcomes). While we continue to build upon the models to better characterize weak toxicants, there is current utility in saving resources by being able to screen out strong reproductive toxicants before investing in vivo testing. The modeling approach (machine learning models) is offered for assessing the rat multigeneration reproductive toxicity of chemicals. Our results suggest that machine learning may be a promising alternative approach to evaluate the potential reproductive toxicity of chemicals.

consensus model↗

Blockchain for Fault-Tolerant Grid Operations

Radial topology and vast geographic coverage make distribution systems prone to widespread power outages upon the failure of a single (or multiple) upstream component. Fault-handling algorithms depend heavily on correct estimations of the system’s state to effectively isolate the affected area and reduce the number of affected customers while maintaining operational safety. The work described here leverages the core features of distributed, consensus-based decision-making processes and the immutability of blockchain, and demonstrates their value in improving fault-tolerant grid operations. In this work, blockchain was used to create a trusted data-sharing platform that enables independent actors to reconstruct the system state; this enables distributed resources to make intelligent decisions with limited knowledge. Although the process requires data sharing, its algorithms have been designed to limit the amount of private information that is exchanged, which helps preserve business-sensitive data and maintain customer privacy. In addition, by reducing the information that must be shared, the communication requirements are also reduced; (however, an in-depth analysis of the communication requirements is beyond the scope of this project). The proposed use cases are intended to represent a foundational basis for third parties to develop functional solutions that can eventually be deployed in the field. To further provide guidance, the envisioned use cases have incorporated design requirements that consider the blockchain characteristics and a need to limit information from surrounding resources, which preserve the assumption and the possibility that such resources could belong to different entities. This report presents a detailed design of the three use cases with the tools needed to enable the analysis being tested. The implemented gross error detection method can detect mismatches when the error exceeds 3.8 times the sensor’s rated accuracy. Detection of the circuit breaker state successfully identified the correct states across all simulation tests. A distribution-system power-flow solution in the simulator OpenDSS generally possesses a convergency tolerance of 0.01% on the voltage magnitude. The evaluation of possible reconnection using voltage magnitude—preserving the data ownership—has a voltage magnitude difference smaller than 0.001% from the OpenDSS result. The results preserving data ownership have a difference within the expected power flow tolerance with full knowledge of the system, which surpasses expectations.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Projected Multi-Agent Consensus Equilibrium (PMACE) With Application to Ptychography

Multi-Agent Consensus Equilibrium (MACE) formulates an inverse imaging problem as a balance among multiple update agents such as data-fitting terms and denoisers. However, each such agent operates on a separate copy of the full image, leading to redundant memory use and slow convergence when each agent affects only a small subset of the full image. In this article, we extend MACE to Projected Multi-Agent Consensus Equilibrium (PMACE), in which each agent updates only a projected component of the full image, thus greatly reducing memory use for some applications. We describe PMACE in terms of an equilibrium problem and an equivalent fixed-point problem and show that in most cases the PMACE equilibrium is not the solution of an optimization problem. To demonstrate the value of PMACE, we apply it to the problem of ptychography, in which a sample is reconstructed from the diffraction patterns resulting from coherent X-ray illumination at multiple overlapping spots. In our PMACE formulation, each spot corresponds to a separate data-fitting agent, with the final solution found as an equilibrium among all the agents. In conclusion, our results demonstrate that the PMACE reconstruction algorithm generates more accurate reconstructions at a lower computational cost than existing ptychography algorithms when the spots are sparsely sampled.

97 MATHEMATICS AND COMPUTING↗

Optimizing Cell-based Antimicrobials through Pooled Genomic Libraries

DNA synthesis and assembly technologies ushered in through synthetic biology have great promise for biomanufacturing, bioremediation, and the development of living therapeutics. Unfortunately, predicting sequence to function relationships, including for biosynthetic pathways expressed in a new host organism, is difficult and often requires many iterative cycles of design, construction, and testing. We are working to develop data-driven approaches to identify the genetic determinants of growth defects and productivity for the expression of a cell-based antimicrobial. We assayed the growth, pigment production, and antimicrobial activity of a collection of over 10,000 genetic mutants of the violacein biosynthetic pathway and sequenced the genetic variation of these mutants. Through this project, we have developed an innovative codebase to automate the determination of pigmentation and antimicrobial clearing diameter for tens of thousands of genetic mutants cultivated on agar dishes. Further, we have written DNA sequence analysis code to demultiplex & provide consensus sequences from high-throughput PacBio long-read circular consensus sequencing (CCS) datasets. From this foundation, we plan to map DNA sequence to function to predict an optimal genetic design to maximize antimicrobial activity while minimizing deleterious growth effects. The workflows and algorithms developed through this project can be broadly applied to other engineered functions in microbes, uncovering sequence to function relationships for complex phenotypes where function impacts fitness.

59 BASIC BIOLOGICAL SCIENCES↗

Benefits and Limits of Phasing Alleles for Network Inference of Allopolyploid Complexes

Abstract Accurately reconstructing the reticulate histories of polyploids remains a central challenge for understanding plant evolution. Although phylogenetic networks can provide insights into relationships among polyploid lineages, inferring networks may be hindered by the complexities of homology determination in polyploid taxa. We use simulations to show that phasing alleles from allopolyploid individuals can improve phylogenetic network inference under the multispecies coalescent by obtaining the true network with fewer loci compared with haplotype consensus sequences or sequences with heterozygous bases represented as ambiguity codes. Phased allelic data can also improve divergence time estimates for networks, which is helpful for evaluating allopolyploid speciation hypotheses and proposing mechanisms of speciation. To achieve these outcomes in empirical data, we present a novel pipeline that leverages a recently developed phasing algorithm to reliably phase alleles from polyploids. This pipeline is especially appropriate for target enrichment data, where the depth of coverage is typically high enough to phase entire loci. We provide an empirical example in the North American Dryopteris fern complex that demonstrates insights from phased data as well as the challenges of network inference. We establish that our pipeline (PATÉ: Phased Alleles from Target Enrichment data) is capable of recovering a high proportion of phased loci from both diploids and polyploids. These data may improve network estimates compared with using haplotype consensus assemblies by accurately inferring the direction of gene flow, but statistical nonidentifiability of phylogenetic networks poses a barrier to inferring the evolutionary history of reticulate complexes.

Evolutionary Biology↗

Cooperative Systems in Presence of Cyber-Attacks: A Unified Framework for Resilient Control and Attack Identification

Here, this paper considers a cooperative control problem in presence of unknown attacks. The attacker aims at destabilizing the consensus dynamics by intercepting the system’s communication network and corrupting its local state feedback. We first revisit the virtual network based resilient control proposed in our previous work and provide a new interpretation and insights into its implementation. Based on these insights, a novel distributed algorithm is presented to detect and identify the compromised communication links. It is shown that it is not possible for the adversary to launch a harmful and stealthy attack by only manipulating the physical states being exchanged via the network. In addition, a new virtual network is proposed which makes it more difficult for the adversary to launch a stealthy attack even though it is also able to manipulate information being exchanged via the virtual network. A numerical example demonstrates that the proposed control framework achieves simultaneously resilient operation and real-time attack identification.

97 MATHEMATICS AND COMPUTING↗

Rapid Computational Identification of Therapeutic Targets for Pathogens

Biological threats continue to persist and evolve as an important challenge to national security. There are multiple ways in which novel viral pathogens could emerge to pose a serious threat to human health. This project developed a pathogen target identification tool that can rapidly respond to a novel or emerging viral biological threat. A set of computational tools were developed that provide detailed information on the newly sequenced genes, their protein products and the drug target sites for the proteins that are best suited for biological countermeasure development. Three key innovations were developed in the project. 1) Development of a new extensive database of protein pocket structures with structure-based search algorithms to rapidly link novel protein targets with the complete collection of previously experimentally solved protein structures. 2) A novel clustering pipeline was introduced to group matching structures and associated small-molecule binding ligands into a consensus protein pocket with the associated small-molecule chemotypes predicted to fit in the pocket site. The matching experimentally solved structures were used to inform the value of different target sites. 3) Where there are viral protein targets with pockets structurally matched to similar human proteins, a biological knowledge graph, which links molecular interactions with human disease, was used to further assess the potential negative impact of a viral protein target with similarities to human proteins that could have important off target side effects. In total, the project produced a new resource for rapid and detailed assessment of promising targets for countermeasures, reflecting the ongoing wet lab, clinical, and computational data being collected. These capabilities will improve the ability to respond to a biological threat in multiple domains.

59 BASIC BIOLOGICAL SCIENCES↗

A Blockchain and PKI-Based Secure Vehicle-to-Vehicle Energy-Trading Protocol

With the increasing awareness for sustainable future and green energy, the demand for electric vehicles (EVs) is growing rapidly, thus placing immense pressure on the energy grid. To alleviate this, local trading between EVs should be encouraged. In this paper, we propose a blockchain and public key infrastructure (PKI)-based secure vehicle-to-vehicle (V2V) energy-trading protocol. A permissioned blockchain utilizing the proof of authority (PoA) consensus and smart contracts is used to securely store data. Encrypted communication is ensured through transport layer security (TLS), with PKI managing the necessary digital certificates and keys. A multi-leader, multi-follower Stackelberg game-based trade algorithm is formulated to determine the optimal energy demands, supplies, and prices. Finally, we propose a detailed communication protocol that ties all the components together, enabling smooth interaction between them. Key findings, such as system behavior and performance, scalability of the trade algorithm and the blockchain, smart contract execution costs, etc., are presented through numerical results by implementing and simulating the protocol in various scenarios. This work not only enhances local energy trading among EVs, encouraging efficient energy usage and reducing burden on the power grid, but also paves a way for future research in sustainable energy management.

Stackelberg game↗

The DESI DR1 peculiar velocity survey: growth rate measurements from the maximum likelihood fields method

We present the constraint on the growth rate of structure from the combination of DESI DR1 BGS sample, Fundamental Plane, and Tully-Fisher peculiar velocity catalogues using the maximum likelihood fields method. The combined catalogue contains 415,523 galaxy redshifts and 76,616 peculiar velocity measurements. To handle the large amount of data in the DESI DR1 peculiar velocity catalogue, we significantly improve the computational efficiency by rewriting the algorithm with JAX. After removing outliers and Tully-Fisher galaxies that are affected by systematics, we find fσ 8 = 0.483 -0.043 +0.080 (stat) ± 0.018(sys), consistent within 1σ with the power spectrum and correlation function analysis using the same dataset. Combining all three measurements with appropriate correlations, the consensus measurement is fσ 8 (z eff = 0.07) = 0.450±0.055, consistent with Planck +ΛCDM cosmology (fσ 8 = 0.449±0.008). Combining with the high redshift growth rate of structure measurements from DESI ShapeFit, the constraint on the growth index is γ = 0.58±0.11, consistent with GR.

cosmic flows↗

Graph Analytics on Jellyfish topology

Because large unstructured datasets is important for many science domains, distributed graph analytics is critical to many scientists. Unfortunately, obtaining scaling and performance for irregular communication is challenging because contemporary network interconnects are primarily designed to maximize bandwidths of fixed-neighborhoods large-message exchanges (e.g., stencils). Although there is no consensus on the “best” network topologies for irregular communication, unstructured graph-based interconnects can be more suitable. We analyze three popular graph workloads – clustering, pattern enumeration, and traversal — on comparable networks (in terms of resources and costs) constructed from Jellyfish Random Regular, Dragonfly and Fat tree topologies, varying the routing algorithms. Using packet-level simulations, we demonstrate up to 60% improvement in communication time with Jellyfish due to diversity of the short paths between arbitrary endpoints, which can reduce overall network stalls and congestion.

Graph Analytics, network topology, interconnect, H↗

Robust Decentralized Learning Using ADMM With Unreliable Agents

Many signal processing and machine learning problems can be formulated as consensus optimization problems which can be solved efficiently via a cooperative multi-agent system. However, the agents in the system can be unreliable due to a variety of reasons: noise, faults and attacks. Providing erroneous updates leads the optimization process in a wrong direction, and degrades the performance of distributed machine learning algorithms. This paper considers the problem of decentralized learning using ADMM in the presence of unreliable agents. First, we rigorously analyze the effect of erroneous updates (in ADMM learning iterations) on the convergence behavior of the multi-agent system. We show that the algorithm linearly converges to a neighborhood of the optimal solution under certain conditions and characterize the neighborhood size analytically. Next, we provide guidelines for network design to achieve a faster convergence to the neighborhood. Here, we also provide conditions on the erroneous updates for exact convergence to the optimal solution. Finally, to mitigate the influence of unreliable agents, we propose ROAD , a robust variant of ADMM, and show its resilience to unreliable agents with an exact convergence to the optimum.

97 MATHEMATICS AND COMPUTING↗

Integrating multimodal data through interpretable heterogeneous ensembles

Motivation: Integrating multimodal data represents an effective approach to predicting biomedical characteristics, such as protein functions and disease outcomes. However, existing data integration approaches do not sufficiently address the heterogeneous semantics of multimodal data. In particular, early and intermediate approaches that rely on a uniform integrated representation reinforce the consensus among the modalities but may lose exclusive local information. The alternative late integration approach that can address this challenge has not been systematically studied for biomedical problems. Results: We propose Ensemble Integration (EI) as a novel systematic implementation of the late integration approach. EI infers local predictive models from the individual data modalities using appropriate algorithms and uses heterogeneous ensemble algorithms to integrate these local models into a global predictive model. We also propose a novel interpretation method for EI models. We tested EI on the problems of predicting protein function from multimodal STRING data and mortality due to coronavirus disease 2019 (COVID-19) from multimodal data in electronic health records. We found that EI accomplished its goal of producing significantly more accurate predictions than each individual modality. It also performed better than several established early integration methods for each of these problems. The interpretation of a representative EI model for COVID-19 mortality prediction identified several disease-relevant features, such as laboratory test (blood urea nitrogen and calcium) and vital sign measurements (minimum oxygen saturation) and demographics (age). These results demonstrated the effectiveness of the EI framework for biomedical data integration and predictive modeling.

59 BASIC BIOLOGICAL SCIENCES↗

Fidelity-preserving enhancement of ptychography with foundational text-to-image models

Ptychographic phase retrieval enables high-resolution imaging of complex samples but often suffers from artifacts such as grid pathology and multislice crosstalk, which degrade reconstructed images. We propose a plug-and-play (PnP) framework that integrates physics model-based phase retrieval with text-guided image editing using foundational diffusion models. By employing the alternating direction method of multipliers, our approach ensures consensus between data fidelity and artifact removal subproblems, maintaining physical consistency while enhancing image quality. Artifact removal is achieved using a text-guided diffusion image editing method (LEDITS++) with a pre-trained foundational diffusion model, allowing users to specify artifacts for removal in natural language. Demonstrations on simulated and experimental datasets show significant improvements in artifact suppression and structural fidelity, validated by metrics such as peak signal-to-noise ratio and diffraction pattern consistency. This work highlights the combination of text-guided generative models and model-based phase retrieval algorithms as a transferable and fidelity-preserving method for high-quality diffraction imaging.

image editing↗

Key predictors of soil organic matter vulnerability to mineralization differ with depth at a continental scale

Abstract Soil organic matter (SOM) is the largest terrestrial pool of organic carbon, and potential carbon-climate feedbacks involving SOM decomposition could exacerbate anthropogenic climate change. However, our understanding of the controls on SOM mineralization is still incomplete, and as such, our ability to predict carbon-climate feedbacks is limited. To improve our understanding of controls on SOM decomposition, A and upper B horizon soil samples from 26 National Ecological Observatory Network (NEON) sites spanning the conterminous U.S. were incubated for 52 weeks under conditions representing site-specific mean summer temperature and sample-specific field capacity (−33 kPa) water potential. Cumulative carbon dioxide respired was periodically measured and normalized by soil organic C content to calculate cumulative specific respiration (CSR), a metric of SOM vulnerability to mineralization. The Boruta algorithm, a feature selection algorithm, was used to select important predictors of CSR from 159 variables. A diverse suite of predictors was selected (12 for A horizons, 7 for B horizons) with predictors falling into three categories corresponding to SOM chemistry, reactive Fe and Al phases, and site moisture availability. The relationship between SOM chemistry predictors and CSR was complex, while sites that had greater concentrations of reactive Fe and Al phases or were wetter had lower CSR. Only three predictors were selected for both horizon types, suggesting dominant controls on SOM decomposition differ by horizon. Our findings contribute to the emerging consensus that a broad array of controls regulates SOM decomposition at large scales and highlight the need to consider changing controls with depth.

59 BASIC BIOLOGICAL SCIENCES↗

Repetitive proteins that undergo large conformational changes evade structural prediction algorithms

Protein structure prediction algorithms, such as AlphaFold, have accelerated protein design and advanced the understanding of the relationship between amino acid sequence and protein structure. However, these algorithms are limited in their ability to predict the structures of conformationally dynamic, intrinsically disordered, and stimuli-responsive proteins. To evaluate sequence-to-structure predictions of such challenging proteins, we explored a class of conformationally dynamic, repeats-in-toxin (RTX) proteins. RTX proteins adopt intrinsically disordered conformations in the absence of calcium and undergo reversible folding into β-roll structures upon binding to calcium. RTX proteins are characterized by tandem repeats of the sequence GGXGXDXUX, in which X can be any amino acid and U is an aliphatic amino acid. We designed RTX sequence variants with global substitutions of nonconserved amino acids, tandem repeats of consensus sequences GGAGXDTLY, and tandem repeats of scrambled sequences GGAGXDTYL. AlphaFold2 and AlphaFold3 predicted that all of these RTX variants adopt β-roll structures, characteristic of wild-type RTX bound to calcium. However, modeling the predicted structures with molecular dynamics simulations and characterizing the protein variants with circular dichroism spectroscopy, small-angle x-ray scattering, and x-ray crystallography revealed that variants adopt diverse, sequence-dependent structures in the absence and presence of calcium. To better design proteins for applications in biotechnology and sustainability, it is critical to build predictive tools that consider intrinsically disordered protein states and validate these tools with multi-mode, multi-scale experimental data.

Chang, Marina P. [Stanford Univ., CA (United State↗