Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Supplementary Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

2013 Metropolitan Area Planning Agency External Travel Survey

# 2013 Metropolitan Area Planning Agency External Travel Survey The 2013 Metropolitan Area Planning Agency (MAPA) External Travel Survey was conducted to measure and identify travel patterns into, within, and out of the greater Omaha/Council Bluffs metropolitan area in Nebraska. MAPA sponsored the survey in conjunction with the Federal Highway Administration, the Nebraska Department of Roads, and the Iowa Department of Transportation. ## Data Collection Agency MAPA conducted the survey. ## Methodology The purpose of the survey was to collect information and data needed as input for MAPA’s travel-demand model. The survey employed a combination of nine survey methods and data-collection activities, including Bluetooth technology, intercept surveys, postcard handouts, travel-time studies, vehicle classification counts, and a web-based survey. ## Survey Records Survey records include a total of 729 participants. ## Transportation Data This study provides Bluetooth records and supplementary data for 17,434 passenger trips and 3,123 commercial trips, accounting for 714,218 vehicle miles traveled. Transportation data are available as zipped files. [Download Winzip](http://www.winzip.com/downwz.htm).

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

2013 Metropolitan Area Planning Agency External Travel Survey

# 2013 Metropolitan Area Planning Agency External Travel Survey The 2013 Metropolitan Area Planning Agency (MAPA) External Travel Survey was conducted to measure and identify travel patterns into, within, and out of the greater Omaha/Council Bluffs metropolitan area in Nebraska. MAPA sponsored the survey in conjunction with the Federal Highway Administration, the Nebraska Department of Roads, and the Iowa Department of Transportation. ## Data Collection Agency MAPA conducted the survey. ## Methodology The purpose of the survey was to collect information and data needed as input for MAPA’s travel-demand model. The survey employed a combination of nine survey methods and data-collection activities, including Bluetooth technology, intercept surveys, postcard handouts, travel-time studies, vehicle classification counts, and a web-based survey. ## Survey Records Survey records include a total of 729 participants. ## Transportation Data This study provides Bluetooth records and supplementary data for 17,434 passenger trips and 3,123 commercial trips, accounting for 714,218 vehicle miles traveled. Transportation data are available as zipped files. [Download Winzip](http://www.winzip.com/downwz.htm).

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

2013 Metropolitan Area Planning Agency External Travel Survey

# 2013 Metropolitan Area Planning Agency External Travel Survey The 2013 Metropolitan Area Planning Agency (MAPA) External Travel Survey was conducted to measure and identify travel patterns into, within, and out of the greater Omaha/Council Bluffs metropolitan area in Nebraska. MAPA sponsored the survey in conjunction with the Federal Highway Administration, the Nebraska Department of Roads, and the Iowa Department of Transportation. ## Data Collection Agency MAPA conducted the survey. ## Methodology The purpose of the survey was to collect information and data needed as input for MAPA’s travel-demand model. The survey employed a combination of nine survey methods and data-collection activities, including Bluetooth technology, intercept surveys, postcard handouts, travel-time studies, vehicle classification counts, and a web-based survey. ## Survey Records Survey records include a total of 729 participants. ## Transportation Data This study provides Bluetooth records and supplementary data for 17,434 passenger trips and 3,123 commercial trips, accounting for 714,218 vehicle miles traveled. Transportation data are available as zipped files. [Download Winzip](http://www.winzip.com/downwz.htm).

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

2013 Metropolitan Area Planning Agency External Travel Survey

# 2013 Metropolitan Area Planning Agency External Travel Survey The 2013 Metropolitan Area Planning Agency (MAPA) External Travel Survey was conducted to measure and identify travel patterns into, within, and out of the greater Omaha/Council Bluffs metropolitan area in Nebraska. MAPA sponsored the survey in conjunction with the Federal Highway Administration, the Nebraska Department of Roads, and the Iowa Department of Transportation. ## Data Collection Agency MAPA conducted the survey. ## Methodology The purpose of the survey was to collect information and data needed as input for MAPA’s travel-demand model. The survey employed a combination of nine survey methods and data-collection activities, including Bluetooth technology, intercept surveys, postcard handouts, travel-time studies, vehicle classification counts, and a web-based survey. ## Survey Records Survey records include a total of 729 participants. ## Transportation Data This study provides Bluetooth records and supplementary data for 17,434 passenger trips and 3,123 commercial trips, accounting for 714,218 vehicle miles traveled. Transportation data are available as zipped files. [Download Winzip](http://www.winzip.com/downwz.htm).

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

RCSB Protein Data Bank 1D tools and services

Abstract Motivation Interoperability between polymer sequences and structural data is essential for providing a complete picture of protein and gene features and helping to understand biomolecular function. Results Herein, we present two resources designed to improve interoperability between the RCSB Protein Data Bank, the NCBI and the UniProtKB data resources and visualize integrated data therefrom. The underlying tools provide a flexible means of mapping between the different coordinate spaces and an interactive tool allows convenient visualization of the 1-dimensional data over the web. Availabilityand implementation https://1d-coordinates.rcsb.org and https://rcsb.github.io/rcsb-saguaro. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

Do-calculus enables estimation of causal effects in partially observed biomolecular pathways

Abstract Motivation Estimating causal queries, such as changes in protein abundance in response to a perturbation, is a fundamental task in the analysis of biomolecular pathways. The estimation requires experimental measurements on the pathway components. However, in practice many pathway components are left unobserved (latent) because they are either unknown, or difficult to measure. Latent variable models (LVMs) are well-suited for such estimation. Unfortunately, LVM-based estimation of causal queries can be inaccurate when parameters of the latent variables are not uniquely identified, or when the number of latent variables is misspecified. This has limited the use of LVMs for causal inference in biomolecular pathways. Results In this article, we propose a general and practical approach for LVM-based estimation of causal queries. We prove that, despite the challenges above, LVM-based estimators of causal queries are accurate if the queries are identifiable according to Pearl’s do-calculus and describe an algorithm for its estimation. We illustrate the breadth and the practical utility of this approach for estimating causal queries in four synthetic and two experimental case studies, where structures of biomolecular pathways challenge the existing methods for causal query estimation. Availability and implementation The code and the data documenting all the case studies are available at https://github.com/srtaheri/LVMwithDoCalculus. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

A Geometry-Driven Longitudal Topic Model

A simple and scalable framework for longitudinal analysis of Twitter data is developed that combines latent topic models with computational geometric methods. Dimensionality reduction tools from computational geometry are applied to learn the intrinsic manifold on which the latent, temporal topics reside. Then shortest path distances on the manifold are used to link together these topics. The proposed framework permits visualization of the low-dimensional embedding which provides clear interpretation of the complex, high-dimensional trajectories that may exist among latent topics. Practical application of the proposed framework is demonstrated through its ability to capture and effectively visualize natural progression of latent COVID-19 related topics learned from Twitter data. Interpretability of the trajectories is achieved by comparing to real-world events. In addition, the framework permits study of spatial variation in Twitter behavior for learned topics. The analysis demonstrates that the proposed framework is able to capture granular-level impact of COVID-19 on public discussions. We end by arguing that Twitter data, when analyzed within the proposed framework, can serve as a valuable supplementary data stream for COVID-related studies.

97 MATHEMATICS AND COMPUTING↗

Complete collision data set for electrons scattering on molecular hydrogen and its isotopologues: III. Vibrational excitation via electronic excitation and radiative decay

We present state-resolved cross sections for vibrational excitation via electronic excitation followed by radiative decay (ERD), for electrons scattering on all bound vibrational levels of the ground electronic state ($X\hspace{1.5 mm} ^1Σ_g^+$) of molecular hydrogen and its isotopologues (H 2 , HD, HT, D 2 , DT and T 2 ). We consider excitation of the singlet $\textit{n}$ = 2–3 states (where n refers to the united-atoms-limit principle quantum number) and account for all possible decay pathways back to the bound vibrational levels of the ground electronic state ($X\hspace{1.5 mm} ^1Σ_g^+$) to produce an estimate for the ERD cross sections. A selection of the results are presented in graphical form and a full data set is provided as both numerical values and analytic fit functions in supplementary data files. The uncertainty in the cross sections is estimated to be 11%, except for scattering on the highest bound vibrational level of HD, HT, D 2 , and DT, where the uncertainty is 21%. The data can be downloaded from the MCCC database at mccc-db.org.

74 ATOMIC AND MOLECULAR PHYSICS↗

CONSTAX2: improved taxonomic classification of environmental DNA markers

Abstract Summary CONSTAX—the CONSensus TAXonomy classifier—was developed for accurate and reproducible taxonomic annotation of fungal rDNA amplicon sequences and is based upon a consensus approach of RDP, SINTAX and UTAX algorithms. CONSTAX2 extends these features to classify prokaryotes as well as eukaryotes and incorporates BLAST-based classifiers to reduce classification errors. Additionally, CONSTAX2 implements a conda-installable command-line tool with improved classification metrics, faster training, multithreading support, capacity to incorporate external taxonomic databases and new isolate matching and high-level taxonomy tools, replete with documentation and example tutorials. Availability and implementation CONSTAX2 is available at https://github.com/liberjul/CONSTAXv2, and is packaged for Linux and MacOS from Bioconda with use under the MIT License. A tutorial and documentation are available at https://constax.readthedocs.io/en/latest/. Data and scripts associated with the manuscript are available at https://github.com/liberjul/CONSTAXv2_ms_code. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

Improving deep learning-based protein distance prediction in CASP14

Abstract Motivation Accurate prediction of residue–residue distances is important for protein structure prediction. We developed several protein distance predictors based on a deep learning distance prediction method and blindly tested them in the 14th Critical Assessment of Protein Structure Prediction (CASP14). The prediction method uses deep residual neural networks with the channel-wise attention mechanism to classify the distance between every two residues into multiple distance intervals. The input features for the deep learning method include co-evolutionary features as well as other sequence-based features derived from multiple sequence alignments (MSAs). Three alignment methods are used with multiple protein sequence/profile databases to generate MSAs for input feature generation. Based on different configurations and training strategies of the deep learning method, five MULTICOM distance predictors were created to participate in the CASP14 experiment. Results Benchmarked on 37 hard CASP14 domains, the best performing MULTICOM predictor is ranked 5th out of 30 automated CASP14 distance prediction servers in terms of precision of top L/5 long-range contact predictions [i.e. classifying distances between two residues into two categories: in contact (<8 Angstrom) and not in contact otherwise] and performs better than the best CASP13 distance prediction method. The best performing MULTICOM predictor is also ranked 6th among automated server predictors in classifying inter-residue distances into 10 distance intervals defined by CASP14 according to the precision of distance classification. The results show that the quality and depth of MSAs depend on alignment methods and sequence databases and have a significant impact on the accuracy of distance prediction. Using larger training datasets and multiple complementary features improves prediction accuracy. However, the number of effective sequences in MSAs is only a weak indicator of the quality of MSAs and the accuracy of predicted distance maps. In contrast, there is a strong correlation between the accuracy of contact/distance predictions and the average probability of the predicted contacts, which can therefore be more effectively used to estimate the confidence of distance predictions and select predicted distance maps. Availability and implementation The software package, source code and data of DeepDist2 are freely available at https://github.com/multicom-toolbox/deepdist and https://zenodo.org/record/4712084#.YIIM13VKhQM. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

MS1Connect: a mass spectrometry run similarity measure

Abstract Motivation Interpretation of newly acquired mass spectrometry data can be improved by identifying, from an online repository, previous mass spectrometry runs that resemble the new data. However, this retrieval task requires computing the similarity between an arbitrary pair of mass spectrometry runs. This is particularly challenging for runs acquired using different experimental protocols. Results We propose a method, MS1Connect, that calculates the similarity between a pair of runs by examining only the intact peptide (MS1) scans, and we show evidence that the MS1Connect score is accurate. Specifically, we show that MS1Connect outperforms several baseline methods on the task of predicting the species from which a given proteomics sample originated. In addition, we show that MS1Connect scores are highly correlated with similarities computed from fragment (MS2) scans, even though these data are not used by MS1Connect. Availability and implementation The MS1Connect software is available at https://github.com/bmx8177/MS1Connect. Supplementary information Supplementary data are available at Bioinformatics online.

47 OTHER INSTRUMENTATION↗

Snekmer: a scalable pipeline for protein sequence fingerprinting based on amino acid recoding

Abstract Motivation The vast expansion of sequence data generated from single organisms and microbiomes has precipitated the need for faster and more sensitive methods to assess evolutionary and functional relationships between proteins. Representing proteins as sets of short peptide sequences (kmers) has been used for rapid, accurate classification of proteins into functional categories; however, this approach employs an exact-match methodology and thus may be limited in terms of sensitivity and coverage. We have previously used similarity groupings, based on the chemical properties of amino acids, to form reduced character sets and recode proteins. This amino acid recoding (AAR) approach simplifies the construction of protein representations in the form of kmer vectors, which can link sequences with distant sequence similarity and provide accurate classification of problematic protein families. Results Here, we describe Snekmer, a software tool for recoding proteins into AAR kmer vectors and performing either (i) construction of supervised classification models trained on input protein families or (ii) clustering for de novo determination of protein families. We provide examples of the operation of the tool against a set of nitrogen cycling families originally collected using both standard hidden Markov models and a larger set of proteins from Uniprot and demonstrate that our method accurately differentiates these sequences in both operation modes. Availability and implementation Snekmer is written in Python using Snakemake. Code and data used in this article, along with tutorial notebooks, are available at http://github.com/PNNL-CompBio/Snekmer under an open-source BSD-3 license. Supplementary information Supplementary data are available at Bioinformatics Advances online.

59 BASIC BIOLOGICAL SCIENCES↗

CHESS 2025: Leaf Area Index (LAI) for meadow, shrub, tree, and understory vegetation

This dataset contains Leaf Area Index (LAI) measurements made as part of the Colorado Headwaters Ecological Spectroscopy Study (CHESS) during June and July of 2025. Data were collected in the Upper Gunnison Basin, Colorado, across three study domains: the Upper East River (CRBU), Almont Triangle (ALMO), and the Upper Taylor Basin (UPTA). Field observations of LAI were collected within 72 hours of airborne data collection by the National Ecological Observatory Network’s Aerial Observation Platform (NEON AOP). The NEON AOP collected waveform LiDAR (Light Detection and Ranging) and imaging spectrometer data in 426 spectral bands from the visible to shortwave infrared. LAI measurements were collected using the LICOR LAI-2200C Plant Canopy Analyzer following protocols outlined in the instrument manual (LI-COR 2019). Sampling targeted four distinct vegetation types: meadows, shrubs, trees, and aspen forest understory. We have archived data separately by site type because different field methods were used for each. At meadow sites, measurements were made at the four corners of 1m x 1m plots, with the instrument moving inward toward the center of the plot. At shrub sites, we measured the canopies of individual shrubs. At tree sites, we made measurements within a 10m x 10m subplot centered around a focal tree, with 30 observations taken on a regular grid. At aspen understory sites, we measured overstory trees following the tree protocol and understory herbaceous vegetation following the meadow protocol. All measurements included above-canopy (A) and below-canopy (B) readings, with specific protocols for scattering correction measurements in direct-sun conditions. Data were processed using the R package `rlai` (Worsham 2025). This package includes functions to calculate LAI, gap fraction, apparent clumping factor (Ω), scattering correction, and other canopy metrics. Package contents: Full file descriptions appear in ‘flmd.csv’. Files named according to the convention ‘lai_*_summary_data_cleaned.csv’ contain summary values of LAI, apparent clumping factor (Ωapp), and scattering correction factors for each site. These are the analysis-ready products that most data users will work with. Files named ‘lai_*_metadata_cleaned.csv’ contain additional site-level observations made during field collection. We have also archived intermediate and supplementary data for users who wish to check our processing approach or apply alternative methods. ‘raw_lai_2200C.zip’ contains the raw files as read from the LI-COR instrument, with no processing applied, in TXT format. The zip archive contains subdirectories by site type, which are further subdivided by sampling area. Filenames correspond to the sampling site number. ‘intermediate_results.zip’ contains detailed output from the processing routines, in JSON format. The zip archive contains subdirectories by site type; filenames correspond to the sampling site number. ‘scattering_correction_logs.zip’ contains logfiles from the implementation of Kobayashi et al.'s (2013) scattering correction algorithm. The logfiles report values of several parameters at each iteration of the algorithm, as the model converges toward a stable solution. They are intended for users who want to verify scattering correction performance. The zip archive contains subdirectories by site type; filenames correspond to the sampling site number. ‘spot_checks.csv’ reports LAI and other values for a small number of files processed with LI-COR FV2200 software (LI-COR 2013) using the same control parameters as in our R-based approach. Additional metadata are provided in a data dictionary describing column names and definitions (dd.csv), and in a file-level metadata file (flmd.csv). All zip files can be expanded with common archive utilities. TXT, CSV, and JSON files can be ingested into R or Python computing environments or read in common text editor utilities. Geospatial information: Geospatial data for mapping measurement site locations are in the files CHESS_polygons_lai_UTM.geojson, CHESS_polygons_shrub_UTM.geojson, and CHESS_polygons_meadow_UTM.geojson in the companion geospatial package for the 2025 CHESS campaign, ‘CHESS 2025: Location data for field observations and sampling’ (Henderson et al., 2026). CHESS Project Description: The Colorado Headwaters Ecological Spectroscopy Study (CHESS) comprised a multi-week airborne remote sensing and field observation campaign in the Upper Gunnison Basin, Colorado, conducted in June and July of 2025. Airborne remote sensing was conducted by the National Ecological Observatory Network Airborne Observation Platform (NEON AOP), concurrent with a field campaign run by the Rocky Mountain Biological Laboratory (RMBL), the Lawrence Berkeley National Laboratory (LBNL) and SLAC National Accelerator Laboratory Watershed Function Science Focus Area (SFA), and NASA-JPL (Jet Propulsion Laboratory) Earth Surface Mineral Dust Source Investigation (EMIT) program. Between June 10 and July 18, 2025, the NEON AOP flight team collected high-resolution aerial imaging spectroscopy and Light Detection and Ranging (LiDAR) data over three domains: the Upper East River (CRBU), Almont Triangle (ALMO), and the Upper Taylor Basin (UPTA). In coordination with the flights, a field campaign acquired ground-truth observations, including observations of vegetation composition, foliar traits, forest demography, and subsurface properties in 18 core sampling areas within the domains. Additional surface water observations were taken at over 380 point locations. All CHESS campaign datasets can be found within the CHESS ESS-DIVE data portal: https://data.ess-dive.lbl.gov/portals/chess. Funding Acknowledgement: Field and remote-sensing data acquisition was performed under a grant from the National Aeronautics and Space Administration (80NSSC24K1005). This work was also supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. * Todorov and Worsham are co–first authors.

2018 NEON and 2025 CHESS Campaigns↗

RCSB Protein Data Bank: improved annotation, search and visualization of membrane protein structures archived in the PDB

Abstract Motivation Membrane proteins are encoded by approximately one fifth of human genes but account for more than half of all US FDA approved drug targets. Thanks to new technological advances, the number of membrane proteins archived in the PDB is growing rapidly. However, automatic identification of membrane proteins or inference of membrane location is not a trivial task. Results We present recent improvements to the RCSB Protein Data Bank web portal (RCSB PDB, rcsb.org) that provide a wealth of new membrane protein annotations integrated from four external resources: OPM, PDBTM, MemProtMD and mpstruc. We have substantially enhanced the presentation of data on membrane proteins. The number of membrane proteins with annotations available on rcsb.org was increased by ∼80%. Users can search for these annotations, explore corresponding tree hierarchies, display membrane segments at the 1D amino acid sequence level, and visualize the predicted location of the membrane layer in 3D. Availability and implementation Annotations, search, tree data and visualization are available at our rcsb.org web portal. Membrane visualization is supported by the open-source Mol* viewer (molstar.org and github.com/molstar/molstar). Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

Complete collision data set for electrons scattering on molecular hydrogen and its isotopologues: II. Fully vibrationally-resolved electronic excitation of the isotopologues of H 2 (X 1 $Σ^{+}_{g}$)

Here, we present a comprehensive set of vibrationally-resolved cross sections for electron-impact electronic excitation of the isotopologues of molecular hydrogen (D 2 , T 2 , HD, HT, and DT) initially in the ground electronic state. We apply the adiabatic-nuclei molecular convergent close-coupling (MCCC) method to calculate cross sections from threshold to 500 eV for excitation of all bound vibrational levels and dissociative excitation of the B 1 $Σ^{+}_{g}$, C 1 Π u , EF 1 $Σ^{+}_{g}$, B' 1 $Σ^{+}_{g}$, GK 1 $Σ^{+}_{g}$, I 1 Π g , J 1 Δ g , D 1 Π u , H 1 $Σ^{+}_{g}$, b 1 $Σ^{+}_{g}$, c 3 Π u , a 3 $Σ^{+}_{g}$, e 3 $Σ^{+}_{u}$, d 3 Π u , h 3 $Σ^{+}_{g}$, g 3 $Σ^{+}_{g}$, i 3 Π g , and j 3 Δ g electronic states from all bound vibrational levels of the ground electronic (X 1 $Σ^{+}_{g}$) state. Including the previously-published MCCC e-H2 cross sections the data set contains cross sections for over 60,000 electronic and vibrational transitions. The cross sections are presented in graphical form and provided as both numerical values and analytic fit functions in supplementary data files. The data can also be downloaded from the MCCC database at mccc-db.org.

74 ATOMIC AND MOLECULAR PHYSICS↗

ORT: a workflow linking genome-scale metabolic models with reactive transport codes

Abstract Motivation Nutrient and contaminant behavior in the subsurface are governed by multiple coupled hydrobiogeochemical processes which occur across different temporal and spatial scales. Accurate description of macroscopic system behavior requires accounting for the effects of microscopic and especially microbial processes. Microbial processes mediate precipitation and dissolution and change aqueous geochemistry, all of which impacts macroscopic system behavior. As ‘omics data describing microbial processes is increasingly affordable and available, novel methods for using this data quickly and effectively for improved ecosystem models are needed. Results We propose a workflow (‘Omics to Reactive Transport—ORT) for utilizing metagenomic and environmental data to describe the effect of microbiological processes in macroscopic reactive transport models. This workflow utilizes and couples two open-source software packages: KBase (a software platform for systems biology) and PFLOTRAN (a reactive transport modeling code). We describe the architecture of ORT and demonstrate an implementation using metagenomic and geochemical data from a river system. Our demonstration uses microbiological drivers of nitrification and denitrification to predict nitrogen cycling patterns which agree with those provided with generalized stoichiometries. While our example uses data from a single measurement, our workflow can be applied to spatiotemporal metagenomic datasets to allow for iterative coupling between KBase and PFLOTRAN. Availability and implementation Interactive models available at https://pflotranmodeling.paf.subsurfaceinsights.com/pflotran-simple-model/. Microbiological data available at NCBI via BioProject ID PRJNA576070. ORT Python code available at https://github.com/subsurfaceinsights/ort-kbase-to-pflotran. KBase narrative available at https://narrative.kbase.us/narrative/71260 or static narrative (no login required) at https://kbase.us/n/71260/258. Supplementary information Supplementary data are available at Bioinformatics online.

54 ENVIRONMENTAL SCIENCES↗

BioADAPT-MRC: adversarial learning-based domain adaptation improves biomedical machine reading comprehension task

ABSTRACT Motivation Biomedical machine reading comprehension (biomedical-MRC) aims to comprehend complex biomedical narratives and assist healthcare professionals in retrieving information from them. The high performance of modern neural network-based MRC systems depends on high-quality, large-scale, human-annotated training datasets. In the biomedical domain, a crucial challenge in creating such datasets is the requirement for domain knowledge, inducing the scarcity of labeled data and the need for transfer learning from the labeled general-purpose (source) domain to the biomedical (target) domain. However, there is a discrepancy in marginal distributions between the general-purpose and biomedical domains due to the variances in topics. Therefore, direct-transferring of learned representations from a model trained on a general-purpose domain to the biomedical domain can hurt the model’s performance. Results We present an adversarial learning-based domain adaptation framework for the biomedical machine reading comprehension task (BioADAPT-MRC), a neural network-based method to address the discrepancies in the marginal distributions between the general and biomedical domain datasets. BioADAPT-MRC relaxes the need for generating pseudo labels for training a well-performing biomedical-MRC model. We extensively evaluate the performance of BioADAPT-MRC by comparing it with the best existing methods on three widely used benchmark biomedical-MRC datasets—BioASQ-7b, BioASQ-8b and BioASQ-9b. Our results suggest that without using any synthetic or human-annotated data from the biomedical domain, BioADAPT-MRC can achieve state-of-the-art performance on these datasets. Availability and implementation BioADAPT-MRC is freely available as an open-source project at https://github.com/mmahbub/BioADAPT-MRC. Supplementary information Supplementary data are available at Bioinformatics online.

60 APPLIED LIFE SCIENCES↗

MINE 2.0: enhanced biochemical coverage for peak identification in untargeted metabolomics

Abstract Summary Although advances in untargeted metabolomics have made it possible to gather data on thousands of cellular metabolites in parallel, identification of novel metabolites from these datasets remains challenging. To address this need, Metabolic in silico Network Expansions (MINEs) were developed. A MINE is an expansion of known biochemistry which can be used as a list of potential structures for unannotated metabolomics peaks. Here, we present MINE 2.0, which utilizes a new set of biochemical transformation rules that covers 93% of MetaCyc reactions (compared to 25% in MINE 1.0). This results in a 17-fold increase in database size and a 40% increase in MINE database compounds matching unannotated peaks from an untargeted metabolomics dataset. MINE 2.0 is thus a significant improvement to this community resource. Availability and implementation The MINE 2.0 website can be accessed at https://minedatabase.ci.northwestern.edu. The MINE 2.0 web API documentation can be accessed at https://mine-api.readthedocs.io/en/latest/. The data and code underlying this article are available in the MINE-2.0-Paper repository at https://github.com/tyo-nu/MINE-2.0-Paper. MINE 2.0 source code can be accessed at https://github.com/tyo-nu/MINE-Database (MINE construction), https://github.com/tyo-nu/MINE-Server (backend web API) and https://github.com/tyo-nu/MINE-app (web app). Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗