Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Supplementary Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Constraints on the Origin of Mercury’s Large Core from Core-Mantle Differentiation Models

Mercury’s core is notoriously large when compared to other planets of our solar system. The origin of this large core is still uncertain. Available data on the surface composition and internal structure of Mercury from the past MESSENGER mission and future data collected by BepiColombo will continue to provide clues to Mercury’s formation. Here, we present results combining experimental data on elemental distribution between core and mantle with spacecraft data, to estimate bulk Mercury composition. We applied this strategy to major elements (Fe, Si, Mg, Al and O), as well as minor elements (Cr and Ti), and compared derived compositions to chondritic data and the chemical compositions of the other terrestrial planets. Our results show that Mercury has a chemical composition significantly different from all known materials of the solar system. In addition, numerous scenarios were proposed to explain Mercury’s structure, including “chaotic models” such as giant impacts and “orderly models” such as aerodynamic sorting. Here, we tested whether Mercury’s composition can be explained by mantle stripping by impacts. We will show how such a scenario reconciles several features of Mercury’s geochemistry with chondritic data. We will also discuss the outlook of additional constraints from supplementary data potentially collected by BepiColombo.

Mercury↗

Do-calculus enables estimation of causal effects in partially observed biomolecular pathways

Abstract Motivation Estimating causal queries, such as changes in protein abundance in response to a perturbation, is a fundamental task in the analysis of biomolecular pathways. The estimation requires experimental measurements on the pathway components. However, in practice many pathway components are left unobserved (latent) because they are either unknown, or difficult to measure. Latent variable models (LVMs) are well-suited for such estimation. Unfortunately, LVM-based estimation of causal queries can be inaccurate when parameters of the latent variables are not uniquely identified, or when the number of latent variables is misspecified. This has limited the use of LVMs for causal inference in biomolecular pathways. Results In this article, we propose a general and practical approach for LVM-based estimation of causal queries. We prove that, despite the challenges above, LVM-based estimators of causal queries are accurate if the queries are identifiable according to Pearl’s do-calculus and describe an algorithm for its estimation. We illustrate the breadth and the practical utility of this approach for estimating causal queries in four synthetic and two experimental case studies, where structures of biomolecular pathways challenge the existing methods for causal query estimation. Availability and implementation The code and the data documenting all the case studies are available at https://github.com/srtaheri/LVMwithDoCalculus. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

A Geometry-Driven Longitudal Topic Model

A simple and scalable framework for longitudinal analysis of Twitter data is developed that combines latent topic models with computational geometric methods. Dimensionality reduction tools from computational geometry are applied to learn the intrinsic manifold on which the latent, temporal topics reside. Then shortest path distances on the manifold are used to link together these topics. The proposed framework permits visualization of the low-dimensional embedding which provides clear interpretation of the complex, high-dimensional trajectories that may exist among latent topics. Practical application of the proposed framework is demonstrated through its ability to capture and effectively visualize natural progression of latent COVID-19 related topics learned from Twitter data. Interpretability of the trajectories is achieved by comparing to real-world events. In addition, the framework permits study of spatial variation in Twitter behavior for learned topics. The analysis demonstrates that the proposed framework is able to capture granular-level impact of COVID-19 on public discussions. We end by arguing that Twitter data, when analyzed within the proposed framework, can serve as a valuable supplementary data stream for COVID-related studies.

97 MATHEMATICS AND COMPUTING↗

Complete collision data set for electrons scattering on molecular hydrogen and its isotopologues: III. Vibrational excitation via electronic excitation and radiative decay

We present state-resolved cross sections for vibrational excitation via electronic excitation followed by radiative decay (ERD), for electrons scattering on all bound vibrational levels of the ground electronic state ($X\hspace{1.5 mm} ^1Σ_g^+$) of molecular hydrogen and its isotopologues (H 2 , HD, HT, D 2 , DT and T 2 ). We consider excitation of the singlet $\textit{n}$ = 2–3 states (where n refers to the united-atoms-limit principle quantum number) and account for all possible decay pathways back to the bound vibrational levels of the ground electronic state ($X\hspace{1.5 mm} ^1Σ_g^+$) to produce an estimate for the ERD cross sections. A selection of the results are presented in graphical form and a full data set is provided as both numerical values and analytic fit functions in supplementary data files. The uncertainty in the cross sections is estimated to be 11%, except for scattering on the highest bound vibrational level of HD, HT, D 2 , and DT, where the uncertainty is 21%. The data can be downloaded from the MCCC database at mccc-db.org.

74 ATOMIC AND MOLECULAR PHYSICS↗

CONSTAX2: improved taxonomic classification of environmental DNA markers

Abstract Summary CONSTAX—the CONSensus TAXonomy classifier—was developed for accurate and reproducible taxonomic annotation of fungal rDNA amplicon sequences and is based upon a consensus approach of RDP, SINTAX and UTAX algorithms. CONSTAX2 extends these features to classify prokaryotes as well as eukaryotes and incorporates BLAST-based classifiers to reduce classification errors. Additionally, CONSTAX2 implements a conda-installable command-line tool with improved classification metrics, faster training, multithreading support, capacity to incorporate external taxonomic databases and new isolate matching and high-level taxonomy tools, replete with documentation and example tutorials. Availability and implementation CONSTAX2 is available at https://github.com/liberjul/CONSTAXv2, and is packaged for Linux and MacOS from Bioconda with use under the MIT License. A tutorial and documentation are available at https://constax.readthedocs.io/en/latest/. Data and scripts associated with the manuscript are available at https://github.com/liberjul/CONSTAXv2_ms_code. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

Cassini/Huygens Program Archive Plan for Science Data

The purpose of this document is to describe the Cassini/Huygens science data archive system which includes policy, roles and responsibilities, description of science and supplementary data products or data sets, metadata, documentation, software, and archive schedule and methods for archive transfer to the NASA Planetary Data System (PDS).

NASA Planetary Science Data System↗

Improving deep learning-based protein distance prediction in CASP14

Abstract Motivation Accurate prediction of residue–residue distances is important for protein structure prediction. We developed several protein distance predictors based on a deep learning distance prediction method and blindly tested them in the 14th Critical Assessment of Protein Structure Prediction (CASP14). The prediction method uses deep residual neural networks with the channel-wise attention mechanism to classify the distance between every two residues into multiple distance intervals. The input features for the deep learning method include co-evolutionary features as well as other sequence-based features derived from multiple sequence alignments (MSAs). Three alignment methods are used with multiple protein sequence/profile databases to generate MSAs for input feature generation. Based on different configurations and training strategies of the deep learning method, five MULTICOM distance predictors were created to participate in the CASP14 experiment. Results Benchmarked on 37 hard CASP14 domains, the best performing MULTICOM predictor is ranked 5th out of 30 automated CASP14 distance prediction servers in terms of precision of top L/5 long-range contact predictions [i.e. classifying distances between two residues into two categories: in contact (<8 Angstrom) and not in contact otherwise] and performs better than the best CASP13 distance prediction method. The best performing MULTICOM predictor is also ranked 6th among automated server predictors in classifying inter-residue distances into 10 distance intervals defined by CASP14 according to the precision of distance classification. The results show that the quality and depth of MSAs depend on alignment methods and sequence databases and have a significant impact on the accuracy of distance prediction. Using larger training datasets and multiple complementary features improves prediction accuracy. However, the number of effective sequences in MSAs is only a weak indicator of the quality of MSAs and the accuracy of predicted distance maps. In contrast, there is a strong correlation between the accuracy of contact/distance predictions and the average probability of the predicted contacts, which can therefore be more effectively used to estimate the confidence of distance predictions and select predicted distance maps. Availability and implementation The software package, source code and data of DeepDist2 are freely available at https://github.com/multicom-toolbox/deepdist and https://zenodo.org/record/4712084#.YIIM13VKhQM. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

MS1Connect: a mass spectrometry run similarity measure

Abstract Motivation Interpretation of newly acquired mass spectrometry data can be improved by identifying, from an online repository, previous mass spectrometry runs that resemble the new data. However, this retrieval task requires computing the similarity between an arbitrary pair of mass spectrometry runs. This is particularly challenging for runs acquired using different experimental protocols. Results We propose a method, MS1Connect, that calculates the similarity between a pair of runs by examining only the intact peptide (MS1) scans, and we show evidence that the MS1Connect score is accurate. Specifically, we show that MS1Connect outperforms several baseline methods on the task of predicting the species from which a given proteomics sample originated. In addition, we show that MS1Connect scores are highly correlated with similarities computed from fragment (MS2) scans, even though these data are not used by MS1Connect. Availability and implementation The MS1Connect software is available at https://github.com/bmx8177/MS1Connect. Supplementary information Supplementary data are available at Bioinformatics online.

47 OTHER INSTRUMENTATION↗

Global backscatter assessment

The focus of this effort is the development of a global-scale model of aerosol backscatter for laser atmospheric wind sounder (LAWS) design and performance studies. Background parameters are derived from aerosol data sets with global-scale spatial and/or temporal coverage, using objective statistical decomposition and/or a priori stratification based on supplementary data. Backscatter coefficients at the LAWS design wavelength are derived from background aerosol physical, chemical, and optical data, or from direct backscatter measurements at other wavelengths, using background conversion factors. Direct measurements of aerosol backscatter at 10.6 microns from the Royal Signals and Radar Establishment (RSRE) and the Wave Propagation Laboratory (WPL) were selected. The RSRE backscatter data processing code were optimized under low backscatter conditions, performed detailed analyses of collocated intercomparisons between the two lidars, and assisted in the analysis of the long-term backscatter climatologies from the two lidars. Timely presentation of global backscattering experiment (GLOBE) research results to the global geophysical community is required.

Bowdle, David A.↗

Basin-scale estimates of oceanic primary production by remote sensing - The North Atlantic

The monthly averaged CZCS data for 1979 are used to estimate annual primary production at ocean basin scales in the North Atlantic. The principal supplementary data used were 873 vertical profiles of chlorophyll and 248 sets of parameters derived from photosynthesis-light experiments. Four different procedures were tested for calculation of primary production. The spectral model with nonuniform biomass was considered as the benchmark for comparison against the other three models. The less complete models gave results that differed by as much as 50 percent from the benchmark. Vertically uniform models tended to underestimate primary production by about 20 percent compared to the nonuniform models. At horizontal scale, the differences between spectral and nonspectral models were negligible. The linear correlation between biomass and estimated production was poor outside the tropics, suggesting caution against the indiscriminate use of biomass as a proxy variable for primary production.

Platt, Trevor↗

Leeward heat transfer experiments on the shuttle orbiter fuselage

A recent study by Mruk et al. (1975) of convective heat transfer on the leeward surface of a space shuttle orbiter fuselage configuration has yielded experimental data on a fuselage-wing geometry and a bare fuselage (without wings). The present paper summarizes that portion of the investigation related to the bare fuselage tests in which data were obtained for an angle of attack of 90 deg so as to produce nominal planar flow. In addition, some supplementary data were obtained at an angle of attack of 30 deg. The bare fuselage geometry, when immersed in a nearly normal approaching flow, behaves as a cylinder of noncircular cross section and moderate aspect ratio. The data presented were obtained with air as the test gas in a 96-in. hypersonic shock tunnel over a specified range of test conditions. The test program included three values of windward surface temperature and two values of leeward section temperature. Leeward convection is characterized by two values of Stanton number; one for the attached flow just prior to separation and the other for the recirculating zone.

Lamb, J. P.↗

Snekmer: a scalable pipeline for protein sequence fingerprinting based on amino acid recoding

Abstract Motivation The vast expansion of sequence data generated from single organisms and microbiomes has precipitated the need for faster and more sensitive methods to assess evolutionary and functional relationships between proteins. Representing proteins as sets of short peptide sequences (kmers) has been used for rapid, accurate classification of proteins into functional categories; however, this approach employs an exact-match methodology and thus may be limited in terms of sensitivity and coverage. We have previously used similarity groupings, based on the chemical properties of amino acids, to form reduced character sets and recode proteins. This amino acid recoding (AAR) approach simplifies the construction of protein representations in the form of kmer vectors, which can link sequences with distant sequence similarity and provide accurate classification of problematic protein families. Results Here, we describe Snekmer, a software tool for recoding proteins into AAR kmer vectors and performing either (i) construction of supervised classification models trained on input protein families or (ii) clustering for de novo determination of protein families. We provide examples of the operation of the tool against a set of nitrogen cycling families originally collected using both standard hidden Markov models and a larger set of proteins from Uniprot and demonstrate that our method accurately differentiates these sequences in both operation modes. Availability and implementation Snekmer is written in Python using Snakemake. Code and data used in this article, along with tutorial notebooks, are available at http://github.com/PNNL-CompBio/Snekmer under an open-source BSD-3 license. Supplementary information Supplementary data are available at Bioinformatics Advances online.

59 BASIC BIOLOGICAL SCIENCES↗

CHESS 2025: Leaf Area Index (LAI) for meadow, shrub, tree, and understory vegetation

This dataset contains Leaf Area Index (LAI) measurements made as part of the Colorado Headwaters Ecological Spectroscopy Study (CHESS) during June and July of 2025. Data were collected in the Upper Gunnison Basin, Colorado, across three study domains: the Upper East River (CRBU), Almont Triangle (ALMO), and the Upper Taylor Basin (UPTA). Field observations of LAI were collected within 72 hours of airborne data collection by the National Ecological Observatory Network’s Aerial Observation Platform (NEON AOP). The NEON AOP collected waveform LiDAR (Light Detection and Ranging) and imaging spectrometer data in 426 spectral bands from the visible to shortwave infrared. LAI measurements were collected using the LICOR LAI-2200C Plant Canopy Analyzer following protocols outlined in the instrument manual (LI-COR 2019). Sampling targeted four distinct vegetation types: meadows, shrubs, trees, and aspen forest understory. We have archived data separately by site type because different field methods were used for each. At meadow sites, measurements were made at the four corners of 1m x 1m plots, with the instrument moving inward toward the center of the plot. At shrub sites, we measured the canopies of individual shrubs. At tree sites, we made measurements within a 10m x 10m subplot centered around a focal tree, with 30 observations taken on a regular grid. At aspen understory sites, we measured overstory trees following the tree protocol and understory herbaceous vegetation following the meadow protocol. All measurements included above-canopy (A) and below-canopy (B) readings, with specific protocols for scattering correction measurements in direct-sun conditions. Data were processed using the R package `rlai` (Worsham 2025). This package includes functions to calculate LAI, gap fraction, apparent clumping factor (Ω), scattering correction, and other canopy metrics. Package contents: Full file descriptions appear in ‘flmd.csv’. Files named according to the convention ‘lai_*_summary_data_cleaned.csv’ contain summary values of LAI, apparent clumping factor (Ωapp), and scattering correction factors for each site. These are the analysis-ready products that most data users will work with. Files named ‘lai_*_metadata_cleaned.csv’ contain additional site-level observations made during field collection. We have also archived intermediate and supplementary data for users who wish to check our processing approach or apply alternative methods. ‘raw_lai_2200C.zip’ contains the raw files as read from the LI-COR instrument, with no processing applied, in TXT format. The zip archive contains subdirectories by site type, which are further subdivided by sampling area. Filenames correspond to the sampling site number. ‘intermediate_results.zip’ contains detailed output from the processing routines, in JSON format. The zip archive contains subdirectories by site type; filenames correspond to the sampling site number. ‘scattering_correction_logs.zip’ contains logfiles from the implementation of Kobayashi et al.'s (2013) scattering correction algorithm. The logfiles report values of several parameters at each iteration of the algorithm, as the model converges toward a stable solution. They are intended for users who want to verify scattering correction performance. The zip archive contains subdirectories by site type; filenames correspond to the sampling site number. ‘spot_checks.csv’ reports LAI and other values for a small number of files processed with LI-COR FV2200 software (LI-COR 2013) using the same control parameters as in our R-based approach. Additional metadata are provided in a data dictionary describing column names and definitions (dd.csv), and in a file-level metadata file (flmd.csv). All zip files can be expanded with common archive utilities. TXT, CSV, and JSON files can be ingested into R or Python computing environments or read in common text editor utilities. Geospatial information: Geospatial data for mapping measurement site locations are in the files CHESS_polygons_lai_UTM.geojson, CHESS_polygons_shrub_UTM.geojson, and CHESS_polygons_meadow_UTM.geojson in the companion geospatial package for the 2025 CHESS campaign, ‘CHESS 2025: Location data for field observations and sampling’ (Henderson et al., 2026). CHESS Project Description: The Colorado Headwaters Ecological Spectroscopy Study (CHESS) comprised a multi-week airborne remote sensing and field observation campaign in the Upper Gunnison Basin, Colorado, conducted in June and July of 2025. Airborne remote sensing was conducted by the National Ecological Observatory Network Airborne Observation Platform (NEON AOP), concurrent with a field campaign run by the Rocky Mountain Biological Laboratory (RMBL), the Lawrence Berkeley National Laboratory (LBNL) and SLAC National Accelerator Laboratory Watershed Function Science Focus Area (SFA), and NASA-JPL (Jet Propulsion Laboratory) Earth Surface Mineral Dust Source Investigation (EMIT) program. Between June 10 and July 18, 2025, the NEON AOP flight team collected high-resolution aerial imaging spectroscopy and Light Detection and Ranging (LiDAR) data over three domains: the Upper East River (CRBU), Almont Triangle (ALMO), and the Upper Taylor Basin (UPTA). In coordination with the flights, a field campaign acquired ground-truth observations, including observations of vegetation composition, foliar traits, forest demography, and subsurface properties in 18 core sampling areas within the domains. Additional surface water observations were taken at over 380 point locations. All CHESS campaign datasets can be found within the CHESS ESS-DIVE data portal: https://data.ess-dive.lbl.gov/portals/chess. Funding Acknowledgement: Field and remote-sensing data acquisition was performed under a grant from the National Aeronautics and Space Administration (80NSSC24K1005). This work was also supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. * Todorov and Worsham are co–first authors.

2018 NEON and 2025 CHESS Campaigns↗

RCSB Protein Data Bank: improved annotation, search and visualization of membrane protein structures archived in the PDB

Abstract Motivation Membrane proteins are encoded by approximately one fifth of human genes but account for more than half of all US FDA approved drug targets. Thanks to new technological advances, the number of membrane proteins archived in the PDB is growing rapidly. However, automatic identification of membrane proteins or inference of membrane location is not a trivial task. Results We present recent improvements to the RCSB Protein Data Bank web portal (RCSB PDB, rcsb.org) that provide a wealth of new membrane protein annotations integrated from four external resources: OPM, PDBTM, MemProtMD and mpstruc. We have substantially enhanced the presentation of data on membrane proteins. The number of membrane proteins with annotations available on rcsb.org was increased by ∼80%. Users can search for these annotations, explore corresponding tree hierarchies, display membrane segments at the 1D amino acid sequence level, and visualize the predicted location of the membrane layer in 3D. Availability and implementation Annotations, search, tree data and visualization are available at our rcsb.org web portal. Membrane visualization is supported by the open-source Mol* viewer (molstar.org and github.com/molstar/molstar). Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

The impact of dynamic data assimilation on the numerical simulations of the QE II cyclone and an analysis of the jet streak influencing the precyclogenetic environment

A mesoscale numerical model is combined with a dynamic data assimilation via Newtonian relaxation, or 'nudging', to provide initial conditions for subsequent simulations of the QE II cyclone. Both the nudging technique and the inclusion of supplementary data are shown to have a large positive impact on the simulation of the QE II cyclone during the initial phase of rapid cyclone development. Within the initial development period (from 1200 to 1800 UTC 9 September 1978), the dynamic assimilation of operational and bogus data yields a coherent two-layer divergence pattern that is not well defined in the model run using only the operational data and static initialization. Diagnostic analysis based on the simulations show that the initial development of the QE II storm between 0000 UTC 9 September and 0000 UTC 10 September was embedded within an indirect circulation of an intense 300-hPa jet streak, was related to baroclinic processes extending throughout a deep portion of the troposphere, and was associated with a classic two-layer mass-divergence profile expected for an extratropical cyclone.

Manobianco, John↗

Complete collision data set for electrons scattering on molecular hydrogen and its isotopologues: II. Fully vibrationally-resolved electronic excitation of the isotopologues of H 2 (X 1 $Σ^{+}_{g}$)

Here, we present a comprehensive set of vibrationally-resolved cross sections for electron-impact electronic excitation of the isotopologues of molecular hydrogen (D 2 , T 2 , HD, HT, and DT) initially in the ground electronic state. We apply the adiabatic-nuclei molecular convergent close-coupling (MCCC) method to calculate cross sections from threshold to 500 eV for excitation of all bound vibrational levels and dissociative excitation of the B 1 $Σ^{+}_{g}$, C 1 Π u , EF 1 $Σ^{+}_{g}$, B' 1 $Σ^{+}_{g}$, GK 1 $Σ^{+}_{g}$, I 1 Π g , J 1 Δ g , D 1 Π u , H 1 $Σ^{+}_{g}$, b 1 $Σ^{+}_{g}$, c 3 Π u , a 3 $Σ^{+}_{g}$, e 3 $Σ^{+}_{u}$, d 3 Π u , h 3 $Σ^{+}_{g}$, g 3 $Σ^{+}_{g}$, i 3 Π g , and j 3 Δ g electronic states from all bound vibrational levels of the ground electronic (X 1 $Σ^{+}_{g}$) state. Including the previously-published MCCC e-H2 cross sections the data set contains cross sections for over 60,000 electronic and vibrational transitions. The cross sections are presented in graphical form and provided as both numerical values and analytic fit functions in supplementary data files. The data can also be downloaded from the MCCC database at mccc-db.org.

74 ATOMIC AND MOLECULAR PHYSICS↗

ORT: a workflow linking genome-scale metabolic models with reactive transport codes

Abstract Motivation Nutrient and contaminant behavior in the subsurface are governed by multiple coupled hydrobiogeochemical processes which occur across different temporal and spatial scales. Accurate description of macroscopic system behavior requires accounting for the effects of microscopic and especially microbial processes. Microbial processes mediate precipitation and dissolution and change aqueous geochemistry, all of which impacts macroscopic system behavior. As ‘omics data describing microbial processes is increasingly affordable and available, novel methods for using this data quickly and effectively for improved ecosystem models are needed. Results We propose a workflow (‘Omics to Reactive Transport—ORT) for utilizing metagenomic and environmental data to describe the effect of microbiological processes in macroscopic reactive transport models. This workflow utilizes and couples two open-source software packages: KBase (a software platform for systems biology) and PFLOTRAN (a reactive transport modeling code). We describe the architecture of ORT and demonstrate an implementation using metagenomic and geochemical data from a river system. Our demonstration uses microbiological drivers of nitrification and denitrification to predict nitrogen cycling patterns which agree with those provided with generalized stoichiometries. While our example uses data from a single measurement, our workflow can be applied to spatiotemporal metagenomic datasets to allow for iterative coupling between KBase and PFLOTRAN. Availability and implementation Interactive models available at https://pflotranmodeling.paf.subsurfaceinsights.com/pflotran-simple-model/. Microbiological data available at NCBI via BioProject ID PRJNA576070. ORT Python code available at https://github.com/subsurfaceinsights/ort-kbase-to-pflotran. KBase narrative available at https://narrative.kbase.us/narrative/71260 or static narrative (no login required) at https://kbase.us/n/71260/258. Supplementary information Supplementary data are available at Bioinformatics online.

54 ENVIRONMENTAL SCIENCES↗

BioADAPT-MRC: adversarial learning-based domain adaptation improves biomedical machine reading comprehension task

ABSTRACT Motivation Biomedical machine reading comprehension (biomedical-MRC) aims to comprehend complex biomedical narratives and assist healthcare professionals in retrieving information from them. The high performance of modern neural network-based MRC systems depends on high-quality, large-scale, human-annotated training datasets. In the biomedical domain, a crucial challenge in creating such datasets is the requirement for domain knowledge, inducing the scarcity of labeled data and the need for transfer learning from the labeled general-purpose (source) domain to the biomedical (target) domain. However, there is a discrepancy in marginal distributions between the general-purpose and biomedical domains due to the variances in topics. Therefore, direct-transferring of learned representations from a model trained on a general-purpose domain to the biomedical domain can hurt the model’s performance. Results We present an adversarial learning-based domain adaptation framework for the biomedical machine reading comprehension task (BioADAPT-MRC), a neural network-based method to address the discrepancies in the marginal distributions between the general and biomedical domain datasets. BioADAPT-MRC relaxes the need for generating pseudo labels for training a well-performing biomedical-MRC model. We extensively evaluate the performance of BioADAPT-MRC by comparing it with the best existing methods on three widely used benchmark biomedical-MRC datasets—BioASQ-7b, BioASQ-8b and BioASQ-9b. Our results suggest that without using any synthetic or human-annotated data from the biomedical domain, BioADAPT-MRC can achieve state-of-the-art performance on these datasets. Availability and implementation BioADAPT-MRC is freely available as an open-source project at https://github.com/mmahbub/BioADAPT-MRC. Supplementary information Supplementary data are available at Bioinformatics online.

60 APPLIED LIFE SCIENCES↗