Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Supplementary Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

A Statistician’s Overview of Physics-Informed Neural Networks for Spatio-Temporal Data

The recent success of deep neural network models with physical constraints (so-called, Physics-Informed Neural Networks, PINNs) has led to renewed interest in the incorporation of mechanistic information in predictive models. Statisticians and others have long been interested in this problem, which has led to several practical and innovative solutions dating back decades. In this overview, we focus on the problem of data-driven prediction and inference of dynamic spatio-temporal processes that include mechanistic information, such as would be available from partial differential equations, with a strong focus on the quantification of uncertainty associated with data, process, and parameters. Here, we give a brief review of several paradigms and focus our attention on Bayesian implementations given they naturally accommodate uncertainty quantification. We then show that it is straight-forward to include the Bayesian PINN (B-PINN) within the Bayesian hierarchical model (BHM) framework that has long been considered for modeling dynamic spatio-temporal processes. Such a BHM-PINN is illustrated via a simulation study in which a latent nonlinear Burgers’ equation PDE governs the dynamics of Poisson distributed spatio-temporal data. Supplementary materials for this article are available online, including a standardized description of the materials available for reproducing the work.

Bayesian↗

2D radial-azimuthal particle-in-cell benchmark for E × B discharges

In this paper we propose a representative simulation test-case of E × B discharges accounting for plasma wall interactions with the presence of both the electron cyclotron drift instability and the modified-two-stream-instability. Seven independently developed particle-in-cell (PIC) codes have simulated this benchmark case, with the same specified conditions. The characteristics of the different codes and computing times are given. Here, results show that both instabilities were captured in a similar fashion and good agreement between the different PIC codes is reported as main plasma parameters were closely related within a 5% interval. The number of macroparticles per cell was also varied and statistical convergence was reached. Detailed outputs are given in the supplementary data, to be used by other similar groups in the perspective of code verification.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Persistent minimal sequences of SARS-CoV-2

Abstract Motivation Severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) has caused more than 14 million cases and more than half million deaths. Given the absence of implemented therapies, new analysis, diagnosis and therapeutics are of great importance. Results Analysis of SARS-CoV-2 genomes from the current outbreak reveals the presence of short persistent DNA/RNA sequences that are absent from the human genome and transcriptome (PmRAWs). For the PmRAWs with length 12, only four exist at the same location in all SARS-CoV-2. At the gene level, we found one PmRAW of size 13 at the Spike glycoprotein coding sequence. This protein is fundamental for binding in human ACE2 and further use as an entry receptor to invade target cells. Applying protein structural prediction, we localized this PmRAW at the surface of the Spike protein, providing a potential targeted vector for diagnostics and therapeutics. In addition, we show a new pattern of relative absent words (RAWs), characterized by the progressive increase of GC content (Guanine and Cytosine) according to the decrease of RAWs length, contrarily to the virus and host genome distributions. New analysis shows the same property during the Ebola virus outbreak. At a computational level, we improved the alignment-free method to identify pathogen-specific signatures in balance with GC measures and removed previous size limitations. Availability and implementation https://github.com/cobilab/eagle. Supplementary information Supplementary data are available at Bioinformatics online.

Pratas, Diogo↗

A public website for the automated assessment and validation of SARS-CoV-2 diagnostic PCR assays

Abstract Summary Polymerase chain reaction-based assays are the current gold standard for detecting and diagnosing SARS-CoV-2. However, as SARS-CoV-2 mutates, we need to constantly assess whether existing PCR-based assays will continue to detect all known viral strains. To enable the continuous monitoring of SARS-CoV-2 assays, we have developed a web-based assay validation algorithm that checks existing PCR-based assays against the ever-expanding genome databases for SARS-CoV-2 using both thermodynamic and edit-distance metrics. The assay-screening results are displayed as a heatmap, showing the number of mismatches between each detection and each SARS-CoV-2 genome sequence. Using a mismatch threshold to define detection failure, assay performance is summarized with the true-positive rate (recall) to simplify assay comparisons. Availability and implementation The assay evaluation website and supporting software are Open Source and freely available at https://covid19.edgebioinformatics.org/#/assayValidation, https://github.com/jgans/thermonucleotide BLAST and https://github.com/LANL-Bioinformatics/assay_validation. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

VPF-Class: taxonomic assignment and host prediction of uncultivated viruses based on viral protein families

Abstract Motivation Two key steps in the analysis of uncultured viruses recovered from metagenomes are the taxonomic classification of the viral sequences and the identification of putative host(s). Both steps rely mainly on the assignment of viral proteins to orthologs in cultivated viruses. Viral Protein Families (VPFs) can be used for the robust identification of new viral sequences in large metagenomics datasets. Despite the importance of VPF information for viral discovery, VPFs have not yet been explored for determining viral taxonomy and host targets. Results In this work, we classified the set of VPFs from the IMG/VR database and developed VPF-Class. VPF-Class is a tool that automates the taxonomic classification and host prediction of viral contigs based on the assignment of their proteins to a set of classified VPFs. Applying VPF-Class on 731K uncultivated virus contigs from the IMG/VR database, we were able to classify 363K contigs at the genus level and predict the host of over 461K contigs. In the RefSeq database, VPF-class reported an accuracy of nearly 100% to classify dsDNA, ssDNA and retroviruses, at the genus level, considering a membership ratio and a confidence score of 0.2. The accuracy in host prediction was 86.4%, also at the genus level, considering a membership ratio of 0.3 and a confidence score of 0.5. And, in the prophages dataset, the accuracy in host prediction was 86% considering a membership ratio of 0.6 and a confidence score of 0.8. Moreover, from the Global Ocean Virome dataset, over 817K viral contigs out of 1 million were classified. Availability and implementation The implementation of VPF-Class can be downloaded from https://github.com/biocom-uib/vpf-tools. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

A variant selection framework for genome graphs

Abstract Motivation Variation graph representations are projected to either replace or supplement conventional single genome references due to their ability to capture population genetic diversity and reduce reference bias. Vast catalogues of genetic variants for many species now exist, and it is natural to ask which among these are crucial to circumvent reference bias during read mapping. Results In this work, we propose a novel mathematical framework for variant selection, by casting it in terms of minimizing variation graph size subject to preserving paths of length α with at most δ differences. This framework leads to a rich set of problems based on the types of variants [e.g. single nucleotide polymorphisms (SNPs), indels or structural variants (SVs)], and whether the goal is to minimize the number of positions at which variants are listed or to minimize the total number of variants listed. We classify the computational complexity of these problems and provide efficient algorithms along with their software implementation when feasible. We empirically evaluate the magnitude of graph reduction achieved in human chromosome variation graphs using multiple α and δ parameter values corresponding to short and long-read resequencing characteristics. When our algorithm is run with parameter settings amenable to long-read mapping (α = 10 kbp, δ = 1000), 99.99% SNPs and 73% SVs can be safely excluded from human chromosome 1 variation graph. The graph size reduction can benefit downstream pan-genome analysis. Availability and implementation https://github.com/AT-CG/VF. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

Quartet-based inference is statistically consistent under the unified duplication-loss-coalescence model

Abstract Motivation The classic multispecies coalescent (MSC) model provides the means for theoretical justification of incomplete lineage sorting-aware species tree inference methods. This has motivated an extensive body of work on phylogenetic methods that are statistically consistent under MSC. One such particularly popular method is ASTRAL, a quartet-based species tree inference method. Novel studies suggest that ASTRAL also performs well when given multi-locus gene trees in simulation studies. Further, Legried et al. recently demonstrated that ASTRAL is statistically consistent under the gene duplication and loss model (GDL). GDL is prevalent in evolutionary histories and is the first core process in the powerful duplication-loss-coalescence evolutionary model (DLCoal) by Rasmussen and Kellis. Results In this work, we prove that ASTRAL is statistically consistent under the general DLCoal model. Therefore, our result supports the empirical evidence from the simulation-based studies. More broadly, we prove that the quartet-based inference approach is statistically consistent under DLCoal. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

AutoCCS: automated collision cross-section calculation software for ion mobility spectrometry–mass spectrometry

Abstract Motivation Ion mobility spectrometry (IMS) separations are increasingly used in conjunction with mass spectrometry (MS) for separation and characterization of ionized molecular species. Information obtained from IMS measurements includes the ion’s collision cross section (CCS), which reflects its size and structure and constitutes a descriptor for distinguishing similar species in mixtures that cannot be separated using conventional approaches. Incorporating CCS into MS-based workflows can improve the specificity and confidence of molecular identification. At present, there is no automated, open-source pipeline for determining CCS of analyte ions in both targeted and untargeted fashion, and intensive user-assisted processing with vendor software and manual evaluation is often required. Results We present AutoCCS, an open-source software to rapidly determine CCS values from IMS-MS measurements. We conducted various IMS experiments in different formats to demonstrate the flexibility of AutoCCS for automated CCS calculation: (i) stepped-field methods for drift tube-based IMS (DTIMS), (ii) single-field methods for DTIMS (supporting two calibration methods: a standard and a new enhanced method) and (iii) linear calibration for Bruker timsTOF and non-linear calibration methods for traveling wave based-IMS in Waters Synapt and Structures for Lossless Ion Manipulations. We demonstrated that AutoCCS offers an accurate and reproducible determination of CCS for both standard and unknown analyte ions in various IMS-MS platforms, IMS-field methods, ionization modes and collision gases, without requiring manual processing. Availability and implementation https://github.com/PNNL-Comp-Mass-Spec/AutoCCS. Supplementary information Supplementary data are available at Bioinformatics online. Demo datasets are publicly available at MassIVE (Dataset ID: MSV000085979).

47 OTHER INSTRUMENTATION↗

Implementation of a practical Markov chain Monte Carlo sampling algorithm in PyBioNetFit

Abstract Summary Bayesian inference in biological modeling commonly relies on Markov chain Monte Carlo (MCMC) sampling of a multidimensional and non-Gaussian posterior distribution that is not analytically tractable. Here, we present the implementation of a practical MCMC method in the open-source software package PyBioNetFit (PyBNF), which is designed to support parameterization of mathematical models for biological systems. The new MCMC method, am, incorporates an adaptive move proposal distribution. For warm starts, sampling can be initiated at a specified location in parameter space and with a multivariate Gaussian proposal distribution defined initially by a specified covariance matrix. Multiple chains can be generated in parallel using a computer cluster. We demonstrate that am can be used to successfully solve real-world Bayesian inference problems, including forecasting of new Coronavirus Disease 2019 case detection with Bayesian quantification of forecast uncertainty. Availability and implementation PyBNF version 1.1.9, the first stable release with am, is available at PyPI and can be installed using the pip package-management system on platforms that have a working installation of Python 3. PyBNF relies on libRoadRunner and BioNetGen for simulations (e.g. numerical integration of ordinary differential equations defined in SBML or BNGL files) and Dask.Distributed for task scheduling on Linux computer clusters. The Python source code can be freely downloaded/cloned from GitHub and used and modified under terms of the BSD-3 license (https://github.com/lanl/pybnf). Online documentation covering installation/usage is available (https://pybnf.readthedocs.io/en/latest/). A tutorial video is available on YouTube (https://www.youtube.com/watch?v=2aRqpqFOiS4&t=63s). Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

A deep dilated convolutional residual network for predicting interchain contacts of protein homodimers

Abstract Motivation Deep learning has revolutionized protein tertiary structure prediction recently. The cutting-edge deep learning methods such as AlphaFold can predict high-accuracy tertiary structures for most individual protein chains. However, the accuracy of predicting quaternary structures of protein complexes consisting of multiple chains is still relatively low due to lack of advanced deep learning methods in the field. Because interchain residue–residue contacts can be used as distance restraints to guide quaternary structure modeling, here we develop a deep dilated convolutional residual network method (DRCon) to predict interchain residue–residue contacts in homodimers from residue–residue co-evolutionary signals derived from multiple sequence alignments of monomers, intrachain residue–residue contacts of monomers extracted from true/predicted tertiary structures or predicted by deep learning, and other sequence and structural features. Results Tested on three homodimer test datasets (Homo_std dataset, DeepHomo dataset and CASP-CAPRI dataset), the precision of DRCon for top L/5 interchain contact predictions (L: length of monomer in a homodimer) is 43.46%, 47.10% and 33.50% respectively at 6 Å contact threshold, which is substantially better than DeepHomo and DNCON2_inter and similar to Glinter. Moreover, our experiments demonstrate that using predicted tertiary structure or intrachain contacts of monomers in the unbound state as input, DRCon still performs well, even though its accuracy is lower than using true tertiary structures in the bound state are used as input. Finally, our case study shows that good interchain contact predictions can be used to build high-accuracy quaternary structure models of homodimers. Availability and implementation The source code of DRCon is available at https://github.com/jianlin-cheng/DRCon. The datasets are available at https://zenodo.org/record/5998532#.YgF70vXMKsB. Supplementary information Supplementary data are available at Bioinformatics online.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

RF-Net 2: fast inference of virus reassortment and hybridization networks

Abstract Motivation A phylogenetic network is a powerful model to represent entangled evolutionary histories with both divergent (speciation) and convergent (e.g. hybridization, reassortment, recombination) evolution. The standard approach to inference of hybridization networks is to (i) reconstruct rooted gene trees and (ii) leverage gene tree discordance for network inference. Recently, we introduced a method called RF-Net for accurate inference of virus reassortment and hybridization networks from input gene trees in the presence of errors commonly found in phylogenetic trees. While RF-Net demonstrated the ability to accurately infer networks with up to four reticulations from erroneous input gene trees, its application was limited by the number of reticulations it could handle in a reasonable amount of time. This limitation is particularly restrictive in the inference of the evolutionary history of segmented RNA viruses such as influenza A virus (IAV), where reassortment is one of the major mechanisms shaping the evolution of these pathogens. Results Here, we expand the functionality of RF-Net that makes it significantly more applicable in practice. Crucially, we introduce a fast extension to RF-Net, called Fast-RF-Net, that can handle large numbers of reticulations without sacrificing accuracy. In addition, we develop automatic stopping criteria to select the appropriate number of reticulations heuristically and implement a feature for RF-Net to output error-corrected input gene trees. We then conduct a comprehensive study of the original method and its novel extensions and confirm their efficacy in practice using extensive simulation and empirical IAV evolutionary analyses. Availability and implementation RF-Net 2 is available at https://github.com/flu-crew/rf-net-2. Supplementary information Supplementary data are available at Bioinformatics online.

Markin, Alexey (ORCID:0000000332809050)↗

SBbadger: biochemical reaction networks with definable degree distributions

Abstract Motivation An essential step in developing computational tools for the inference, optimization and simulation of biochemical reaction networks is gauging tool performance against earlier efforts using an appropriate set of benchmarks. General strategies for the assembly of benchmark models include collection from the literature, creation via subnetwork extraction and de novo generation. However, with respect to biochemical reaction networks, these approaches and their associated tools are either poorly suited to generate models that reflect the wide range of properties found in natural biochemical networks or to do so in numbers that enable rigorous statistical analysis. Results In this work, we present SBbadger, a python-based software tool for the generation of synthetic biochemical reaction or metabolic networks with user-defined degree distributions, multiple available kinetic formalisms and a host of other definable properties. SBbadger thus enables the creation of benchmark model sets that reflect properties of biological systems and generate the kinetics and model structures typically targeted by computational analysis and inference software. Here, we detail the computational and algorithmic workflow of SBbadger, demonstrate its performance under various settings, provide sample outputs and compare it to currently available biochemical reaction network generation software. Availability and implementation SBbadger is implemented in Python and is freely available at https://github.com/sys-bio/SBbadger and via PyPI at https://pypi.org/project/SBbadger/. Documentation can be found at https://SBbadger.readthedocs.io. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

End-to-end learning of multiple sequence alignments with differentiable Smith–Waterman

Abstract Motivation Multiple sequence alignments (MSAs) of homologous sequences contain information on structural and functional constraints and their evolutionary histories. Despite their importance for many downstream tasks, such as structure prediction, MSA generation is often treated as a separate pre-processing step, without any guidance from the application it will be used for. Results Here, we implement a smooth and differentiable version of the Smith–Waterman pairwise alignment algorithm that enables jointly learning an MSA and a downstream machine learning system in an end-to-end fashion. To demonstrate its utility, we introduce SMURF (Smooth Markov Unaligned Random Field), a new method that jointly learns an alignment and the parameters of a Markov Random Field for unsupervised contact prediction. We find that SMURF learns MSAs that mildly improve contact prediction on a diverse set of protein and RNA families. As a proof of concept, we demonstrate that by connecting our differentiable alignment module to AlphaFold2 and maximizing predicted confidence, we can learn MSAs that improve structure predictions over the initial MSAs. Interestingly, the alignments that improve AlphaFold predictions are self-inconsistent and can be viewed as adversarial. This work highlights the potential of differentiable dynamic programming to improve neural network pipelines that rely on an alignment and the potential dangers of optimizing predictions of protein sequences with methods that are not fully understood. Availability and implementation Our code and examples are available at: https://github.com/spetti/SMURF. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

Dwarf AGNs from Optical Variability for the Origins of Seeds (DAVOS): insights from the dark energy survey deep fields

ABSTRACT We present a sample of 706, z < 1.5 active galactic nuclei (AGNs) selected from optical photometric variability in three of the Dark Energy Survey (DES) deep fields (E2, C3, and X3) over an area of 4.64 deg2. We construct light curves using difference imaging aperture photometry for resolved sources and non-difference imaging PSF photometry for unresolved sources, respectively, and characterize the variability significance. Our DES light curves have a mean cadence of 7 d, a 6-yr baseline, and a single-epoch imaging depth of up to g ∼ 24.5. Using spectral energy distribution (SED) fitting, we find 26 out of total 706 variable galaxies are consistent with dwarf galaxies with a reliable stellar mass estimate ($M_{\ast }\lt 10^{9.5}\, {\rm M}_\odot$; median photometric redshift of 0.9). We were able to constrain rapid characteristic variability time-scales (∼ weeks) using the DES light curves in 15 dwarf AGN candidates (a subset of our variable AGN candidates) at a median photometric redshift of 0.4. This rapid variability is consistent with their low black hole (BH) masses. We confirm the low-mass AGN nature of one source with a high S/N optical spectrum. We publish our catalogue, optical light curves, and supplementary data, such as X-ray properties and optical spectra, when available. We measure a variable AGN fraction versus stellar mass and compare to results from a forward model. This work demonstrates the feasibility of optical variability to identify AGNs with lower BH masses in deep fields, which may be more ‘pristine’ analogues of supermassive BH seeds.

79 ASTRONOMY AND ASTROPHYSICS↗

Dose Coefficient Calculation for Use in Dosimetry Assessment of a Fission-Based Weapon

In the event of a fission-based weapon or improvised nuclear device (IND) detonation, dose coefficients can be harnessed to provide dose assessments for defense, emergency preparedness, and consequence management, as well as to prospectively inform the assessment of radiation biomarkers and development of medical prophylaxis countermeasures for defense and homeland security stakeholders and decision-makers. Although dose coefficients have previously been calculated for this group, they would apply specifically to the studied population, the 1945 Japanese cohort, after which their anthropomorphic computational phantoms were modeled. For this reason, applications to other populations may be limited, and instead, an assessment of a more standardized population is desired. We employed a series of computational human phantoms representing international reference individuals: UF/NCI voxel phantom series containing newborn, 1-, 5-, 10-, 15-, and 35-year-old males and females. Irradiation of the phantoms was simulated using the Monte Carlo N-Particle transport code to determine organ dose coefficients under four idealized irradiation geometries at three distances from the detonation hypocenter at Hiroshima and Nagasaki using DS02 free-in-air prompt neutron and photon fluence spectra. Through these simulations, age-specific dose coefficients were determined for individual organs. Various articulated PIMAL stylized phantoms were simulated as well to estimate the effect of body posture on dose coefficients and determine the effect of posture on dosimetric estimation and reconstruction. Results additionally demonstrate that 137 Cs and the Watt fission spectra are not ideal general surrogate sources for fission weapons, which may be considered for experimental testing of medical countermeasures. Supplementary data provided tabulates the compilation of organ dose-rate coefficients in this study.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Wire-arc Additive Manufacturing Benchmark

This is the dataset associated with the 2022 SRP Additive Manufacturing Prediction Challenge, originally hosted on Github at https://github.com/SRP-AM/SRP_AM_Prediction_Challenge. The benchmark was designed for validating prediction for the temperature history, residual stress, and distortion of an additively manufactured metal part with relatively simple geometry. A calibration problem with the same as-built geometry is provided with measured quantities of interest; including temperature histories at selective locations, post-build residual stress at selective locations, and overall distortion measurements. The challenge problem is presented with a different build sequence (i.e. thermal history). In this dataset, we include the actual recorded calibration and challenge measurements, as well as benchmark template files for testing predictions without incorporating the challenge data. Supplementary files around the materials and setup are available for transparency and reproducibility.

Bachus, Nicholas [UC Davis, Davis, CA]↗

MODFLOW6 models used to evaluate potential stresses and hydrologic conditions driving water-level fluctuations in well ER-5-3-2, Frenchman Flat, Southern Nevada

The hydrograph for well ER-5-3-2 in Frenchman Flat, southern Nevada, has previously unexplained water-level fluctuations. Four, three-dimensional, groundwater models (MODFLOW 6) were developed to evaluate potential stresses and hydrologic conditions affecting the well ER-5-3-2 hydrograph. Four model scenarios were developed that simulated: (1) wellbore leakage without recharge, (2) wellbore leakage with recharge, (3) shallow (low transmissivity) and deep (high transmissivity) carbonate rocks, and (4) lateral heterogeneity of carbonate rocks. Input and output files for the four model scenarios are in the model and output directories, respectively. Hydraulic conductivity, specific storage, and wellbore-leakage rates (when simulated) were estimated with parameter estimation (PEST) by minimizing a weighted composite, sum-of-squares objective function. The objective function was informed by measurement and Tikhonov regularization observations. Measurement observations included drawdowns from the constant-rate aquifer test and water-level altitudes measured in well ER-5-3-2 from 2001-2021. Tikhonov regularization informed hydraulic conductivity and specific storage parameters that were insensitive to measurement observations, where homogeneity was the preferred relation. Batch files, executables, and MODFLOW 6, PEST, and post-processing utilities are in the ancillary directory. Supplementary data also are included in the ancillary directory, including site information, high-frequency water-level and aquifer-test data, transmissivity estimates, water-chemistry data, and water-temperature analyses. This USGS data release contains data, analyses, and model files for the simulations and analysis results described in U.S. Geological Survey Scientific Investigations Report (https://doi.org/10.3133/sir20225132).

54 ENVIRONMENTAL SCIENCES↗