Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Environmental data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Surface and Buried Thermal, and RGB Unexploded Ordnance Data Collection

This document provides a description of a data collection campaign of unexploded ordnance (UXOI) set. The dataset captures a controlled UAV imaging campaign designed to support detection of UXO across varied environmental conditions. Data were collected during three campaigns in Norris and Northeast Knoxville, Tennessee, using RGB, and thermal sensors mounted on Parrot UKR. In total, the dataset contains 9925 images, 26 full-motion video, and approximately 81.99 GB of data, collected across late spring/summer conditions, every hour during sunlight, and multiple surface contexts, including tall grass, short grass, gravel, as well as buried in sand, and other gravel mixtures. The collection was designed to capture thermal and visual variability relevant to UXO detection in agricultural land, bare earth, and subsurface. Review of the imagery showed that ordnance was most detectable during periods of changing solar input, especially approximately 10-60 minutes after sunrise, approximately 20-60 minutes after sunset, and 2-3 min after cloud cover interrupted prolonged solar heating. These conditions increased thermal contrast because many ordnance items retained or released heat differently than the surrounding vegetation and ground surface. This dataset provides a useful resource for developing and evaluating airborne UXO detection methods under realistic field conditions. All ordnance used in the study was inert, and thermal behavior may differ from that of live ordnance. In addition, variation in ordnance type, composition, and placement introduced differences in thermal response that should be considered when interpreting results.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Observed Variability in Convective Cell Characteristics and Near-Storm Environments across the Sea- and Bay-Breeze Fronts in Southeast Texas

Abstract During the DOE Atmospheric Radiation Measurement (ARM) Tracking Aerosol Convection Interactions Experiment (TRACER) IOP spanning June–September 2022, two fixed ARM sites and a mobile team concurrently sampled the airmass heterogeneity across sea- and bay-breeze fronts around the greater Houston metropolitan region. Here, we quantify the spatiotemporal variability between maritime (coastal/bay side of breeze fronts) and continental (inland side of breeze fronts) air masses over 15 IOP days characterized by strong sea-breeze forcing. We analyze environmental profile data from 177 radiosondes and use S- and C-band radar data to track and quantify the variability in attributes of more than 2300 shallow and transitioning cells across different air masses. The composite analysis of environmental profiles indicates that during the early afternoon, the sea-breeze maritime air mass exhibits lower convective available potential energy (CAPE) than the bay-breeze maritime air mass. As the sea breeze advances inland with time, CAPE within the maritime air mass exceeds that of the continental air mass to the north of the breeze fronts. In general, maritime cells have a larger mean composite reflectivity and cell widths than continental cells; however, the response varies between shallow and transitioning cells. Mean composite 20-dB Z echo-top heights, however, are similar across air masses for both shallow and transitioning cells. The continental and maritime inflow air mass for transitioning cells has significantly different mean values for mixed-layer entrainment CAPE, lifted condensation level, level of free condensation, boundary layer depth, and diluted equilibrium level. For shallow cells, only total precipitable water shows a significant difference. Significance Statement The greater Houston metropolitan area is a natural laboratory for understanding the individual impacts of background meteorology and aerosols on convective clouds. Due to its proximity to the Gulf Coast and Galveston Bay, the Houston region experiences a diurnal precipitation cycle in the summer, driven by convection triggered from sea- and bay-breeze fronts. These fronts act as a boundary between air masses with distinct thermodynamic and environmental characteristics. Convergence along these fronts and interactions between storm outflow and the fronts facilitate convection initiation in different mesoscale air masses. This study quantifies the heterogeneity among these air masses while investigating their influence on cloud microphysics. We find that the effect of airmass heterogeneity is more pronounced for the bulk microphysical properties in shallow clouds.

Sharma, Milind↗

Data from "A Bayesian Record Linkage Approach to Applications in Tree Demography Using Overlapping LiDAR Scans"

Processed LiDAR data and environmental covariates from 2015 and 2019 LiDAR scans in the Vicinity of Snodgrass Mountain (Western Colorado, USA), in a geographic subset used in primary analysis for the research paper.This package contains LiDAR-derived canopy height maps for 2015 and 2019, crown polygons derived from the height maps using a segmentation algorithm, and environmental covariates supporting the model of forest growth. Source datasets include August 2015 and August 2019 discrete-return LiDAR point clouds collected by Quantum Geospatial for terrain mapping purposes on behalf of the Colorado Hazard Mapping Program and the Colorado Water Conservation Board. Both datasets adhere to the USGS QL2 quality standard. The point cloud data were processed using the R package lidR to generate a canopy height model representing maximum vegetation height above the ground surface, using a pit-free algorithm.This dataset was compiled to assess how spatial patterns of tree growth in montane and subalpine forests are influenced by water and energy availability. Understanding these growth patterns can provide insight into forest dynamics in the Southern Rocky Mountains under changing climatic conditions.This dataset contains .tif, .csv, and .txt files. This dataset additionally includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata; and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.

54 ENVIRONMENTAL SCIENCES↗

2020 Multiscale Microbial Dynamics Modeling Course

The 2020 Multiscale Microbial Dynamics course is adapted from the virtual 2020 Mutliscale Microbial Dynamics Summer School that was hosted by Environmental Molecular Sciences Laboratory (EMSL), a U.S. Department of Energy (DOE) science user facility located on the Pacific Northwest National Laboratory (PNNL) campus, in collaboration with the Joint Genome Institute (JGI) and the DOE Systems Biology Knowledgebase (KBase). The course course covers how to incorporate microbial metagenomic and environmental metabolite data from watershed ecosystems into metabolic and community modeling using computational frameworks, such as KBase and PFLOTRAN. The curriculum includes lectures and software and data analysis tutorials. All materials are freely accessible to the community as part of the 2020 Microbial Dynamics Summer School Organization in KBase.

54 ENVIRONMENTAL SCIENCES↗

2024 OES-Environmental 2024 State of the Science Report, Chapter 8: Marine Renewable Energy Data and Information Systems

As the marine renewable energy (MRE) sector grows, large amounts of environmental and technical data and information are being collected. When these data and information are openly available, they can be used to guide research and development, inform responsible siting and consenting of projects, and increase stakeholder understanding through transparency. For example, quality environmental data collected during the siting, consenting, construction, operation, and decommissioning of MRE projects can all play key roles in better characterizing baseline conditions, developing effective monitoring and mitigation strategies, and retiring environmental risks through data transferability (see Chapter 6). Ensuring that these data and information are easily discoverable and accessible will help the MRE sector make informed decisions and coexist in an increasingly busy ocean environment.

16 TIDAL AND WAVE POWER↗

The IsoGenie database: an interdisciplinary data management solution for ecosystems biology and environmental research

Modern microbial and ecosystem sciences require diverse interdisciplinary teams that are often challenged in “speaking” to one another due to different languages and data product types. Here we introduce the IsoGenie Database, a de novo developed data management and exploration platform, as a solution to this challenge of accurately representing and integrating heterogenous environmental and microbial data across ecosystem scales. The IsoGenieDB is a public and private data infrastructure designed to store and query data generated by the IsoGenie Project, a ~10 year DOE-funded project focused on discovering ecosystem climate feedbacks in a thawing permafrost landscape. The IsoGenieDB provides (i) a platform for IsoGenie Project members to explore the project’s interdisciplinary datasets across scales through the inherent relationships among data entities, (ii) a framework to consolidate and harmonize the datasets needed by the team’s modelers, and (iii) a public venue that leverages the same spatially explicit, disciplinarily integrated data structure to share published datasets. The IsoGenieDB is also being expanded to cover the NASA-funded Archaea to Atmosphere (A2A) project, which scales the findings of IsoGenie to a broader suite of Arctic peatlands, via the umbrella A2A Database (A2A-DB). The IsoGenieDB’s expandability and flexible architecture allow it to serve as an example ecosystems database.

54 ENVIRONMENTAL SCIENCES↗

Regional-scale soil carbon predictions can be enhanced by transferring global-scale soil–environment relationships

Accurate modelling and mapping soil organic carbon are crucial for supporting soil health restoration and climate change mitigation at both regional and global scales. However, regional soil predictions often suffer from data scarcity and high prediction uncertainty. Utilizing a pre-trained global-to-regional soil carbon predictive model can be a potential solution to address this challenge. Despite its promise, how to construct and apply the global-scale model to enhance regional-scale soil carbon mapping remains largely unexplored. Here, we propose the Global Soil Carbon Pre-trained Model (GSoilCPM), a deep-learning-based domain adaptative model, to enhance regional-scale soil carbon predictions. Based on large amount of environmental covariate data and 106,167 soil samples across the globe, we verify our hypothesis of the effectiveness of this 'global-to-regional' modelling strategy. The pre-trained model can be then transferred and fine-tuned to bridge the regional- and global-scale soil–environment relationships. We applied and validated this modelling strategy in four regional-scale study areas, three in the Northern Hemisphere and one in the Southern Hemisphere, each with distinct environmental background. Compared to traditional modelling approaches as a baseline, four case studies all demonstrated significant improvement in prediction accuracy across diverse environments and varying data availabilities. The average percentage improvement across all regions is 10.93% (absolute values decreased by 1.20 g kg−1 averagely) in MAE and 29.04% (absolute values increased by 0.10 averagely) in CCC. The applicability and future horizons of using GSoilCPM were further discussed. We further reveal that regions with fewer soil samples or lower baseline accuracy benefit more from the pre-trained global model. Our findings highlight the advantages of leveraging the generalized knowledge from global models to enhance specifically localized soil modelling, positioning a potential paradigm shift in digital soil mapping, and far-reaching implications for soil monitoring and land management.

Deep learning↗

Vegetation Warming Experiment: Thaw depth and dGPS locations, Utqiagvik, Alaska, 2021

Thaw depth measurements within and around warming chambers, and in ambient plots located on the Barrow Environmental Observatory (BEO), Utqiagvik, Alaska. Measurements were taken at the start and end of chamber deployment, and two intermediate times during the 2021 growth season. dGPS measurements of chamber and ambient plot locations are also included. The files included in this data package are in .csv format, and include 2 data files and 3 metadata files. This data was recorded as part of the Zero Power Warming (ZPW) vegetation warming experiment. Other datasets under the Vegetation Warming Experiment include data for environmental conditions, leaf physiology, leaf traits, and landscape and plot phenocam images. The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic), was a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research. The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska. Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

54 ENVIRONMENTAL SCIENCES↗

Predicting metabolic modules in incomplete bacterial genomes with MetaPathPredict

The reconstruction of complete microbial metabolic pathways using ‘omics data from environmental samples remains challenging. Computational pipelines for pathway reconstruction that utilize machine learning methods to predict the presence or absence of KEGG modules in incomplete genomes are lacking. Here, we present MetaPathPredict, a software tool that incorporates machine learning models to predict the presence of complete KEGG modules within bacterial genomic datasets. Using gene annotation data and information from the KEGG module database, MetaPathPredict employs deep learning models to predict the presence of KEGG modules in a genome. MetaPathPredict can be used as a command line tool or as a Python module, and both options are designed to be run locally or on a compute cluster. Benchmarks show that MetaPathPredict makes robust predictions of KEGG module presence within highly incomplete genomes.

59 BASIC BIOLOGICAL SCIENCES↗

Heavy-Duty Vehicle Activity Updates for MOVES Using NREL Fleet DNA and CE-CERT Data

The U.S. Environmental Protection Agency's (EPA's) Motor Vehicle Emission Simulator (MOVES) is a publicly available tool used by researchers and policymakers to help understand motor vehicle emission sources at a national, county, and project level. Estimates of heavy-duty activity in the most recent version of the model at the time this work was conducted, MOVES2014, was identified as an area in need of improvement. The start activity in MOVES2014 is based on a limited and dated data set. In addition, MOVES2014 relies on drive cycles that represent on-network activity but do not account for idling activity that occurs on off-network roads, such as at a distribution center, while the truck is queuing or during loading and unloading. As a result, MOVES2014 may currently underestimate the number of starts and idle and soak time for heavy-duty trucks in real-world operation. The National Renewable Energy Laboratory (NREL) has previously leveraged its expansive Fleet DNA database of heavy-duty vehicles to idle and start activity for six of the nine heavy-duty vehicle source types of classes in the MOVES model. The data available in Fleet DNA from 416 conventional, diesel-powered vehicles provided activity estimates from more than 120,000 hours of operation throughout 14,682 vehicle days between October 2006 and January 2016. NREL calculated start fraction, starts per day, soak fraction, and idle fraction by hour of the day for each vehicle type, state, and vocation, and provided results in .CSV files that can be translated to MOVES table inputs. The idle and start activity from this initial analysis of Fleet DNA data was used to develop default idle and start data for heavy-duty vehicles in MOVES3. Satisfied with the results from the Fleet DNA data used for MOVES3, the EPA asked NREL to extend this start/soak/idle analysis using additional data from a larger number of vehicles for a potential future update to the MOVES model. Such a data set was achieved from a project led by the University of California at Riverside, College of Engineering, Center for Environmental Research & Technology (CE-CERT) and funded by California Air Resources Board. Specifically, this data set consists of 90 heavy-duty vehicles operated mainly in California, which can be separated into five of the nine heavy-duty vehicle classes in the MOVES model. In addition, the heavy-duty activity database collected by CE-CERT provided activity estimates from more than 44,000 hours of operation throughout 4,724 vehicle days between November 2014 and September 2016. This report details the analysis of the heavy-duty activity database collected from the University of California at Riverside by providing graphical analysis and context for the start, soak, and idle distributions. The comparison of the related results from both the Fleet DNA and CE-CERT data sets are documented as well.

33 ADVANCED PROPULSION SYSTEMS↗

A Review of the Use of Wearables in Indoor Environmental Quality Studies and an Evaluation of Data Accessibility from a Wearable Device

An understanding of indoor environmental quality (IEQ) and its effects on occupant well-being can inform building system design and operation. The use of wearables in field studies to collect subjective and objective health performance indicators (HPIs) from a large number of occupants could deliver important improvements in IEQ. To facilitate the use of wearables in IEQ studies, there is a need to identify which HPIs should be collected and to evaluate data accessibility from these devices. To address this issue, a literature review of previous IEQ studies was conducted to identify relationships between different IEQ factors and HPIs, with a focus on HPIs that were collected using wearables. A preliminary assessment of data accessibility from a selected wearable device (Fitbit Versa 2) was performed and documented. The review suggested the need to further investigate and collect sleep quality parameters, heart rate, stress response, as well as subjective ratings of comfort using wearables. The data accessibility assessment revealed issues related to missing data points and data resolution from the examined device. A set of recommendations is outlined to inform future studies.

60 APPLIED LIFE SCIENCES↗

Historical Data Analysis Supporting the Data Quality Objectives for the INL Site Environmental Soil Monitoring Program

This document represents the initial evaluation and soil monitoring proposed by Battelle Energy Alliance, LLC (BEA) in 2015. The evaluation included analyses of historical soil monitoring data and soil inventories, current emission estimates, and modeled potential deposition/accumulation patterns. The initially proposed monitoring included a 5-year rotation of in-situ gamma measurements augmented by soil sampling with laboratory analyses near each major active and some inactive facilities. It also proposed rotational in-situ gamma measurements and soil sampling at two centrally located onsite air monitoring locations coinciding with sampling at the traditional offsite soil monitoring locations. The chosen alternative includes only physical soil sampling with laboratory analysis and only at the Radioactive Waste Management Complex (RWMC), the two air monitors and the offsite locations as documented in Data Quality Objectives Supporting the Environmental Soil Monitoring Program for the Idaho National Laboratory (INL) Site, INL/EXT-15-34909, Revision 0, February 2016. The data and evaluations in this document are valid for comparisons with future soil data that may be collected in many INL site locations.

54 ENVIRONMENTAL SCIENCES↗

Contextualizing Non-Powered Dam Site Selection for Archimedes Screw Turbines: A Methodology for Responsible Archimedes Screw Turbine Conversion at Existing Dams

Non-powered dams represent 97% of dams in the United States and their energy generation potential has not been fully realized. The use of an Archimedes screw turbine to generate power at non-powered dams offers a dual benefit; producing electricity, and acting as downstream fish passage, helping to reconnect previously separated ecosystems. In this study, we assess the technical, environmental, social, and economic feasibility of generating power at non-powered U.S. dam sites using Archimedes screw turbines by integrating mechanical constraints, social impact metrics, proximity to infrastructure, and environmental sensitivity data. Results account for future precipitation predictions and show, between 2024 and 2050, the number of sites where Archimedes screw turbines are viable decreases by one site, but overall generation capacity increases due to increased flow rates across persisting locations. Our analysis identified 82 non-powered dam sites with a mean generation capacity of 49 kW that meet the mechanical requirements for Archimedes screw turbine technology in 2024. Our analysis presents a framework for considering social, environmental, and economic impacts of specific turbine technologies to convert non-powered dams to generate power.

Archimedes screw turbine↗

Adaptive Sampling for In Situ Cloud Probe (Final Report)

Clouds play a leading role in the Earth's global energy and solar radiation balance and hydrological cycle. Improving cloud models requires detailed information on the cloud microphysical properties, such as droplet size distribution and number density, liquid water content and cloud composition (droplets, ice particles), which can only be provided by aerial in situ measurements. However, for many atmospheric measurement instruments, the lack of flexibility in selecting the operational mode during operation can lead to uncertainties in sampling and measurement characteristics under continuously varying atmospheric conditions. This SBIR project is developing an advanced, compact optical imaging technology for in situ characterization of cloud hydrometeors. The development involves a deep modification of the existing Mesa Photonics’ Cloud Droplet Measurement System (CDMS) in order to implement real-time automatic adaptive sampling based on the acquired in situ data and environmental parameters. The new system, CDMS-2, implements two measurement modes: side-scatter imaging for smaller hydrometeors and direct bright-field-illumination imaging for larger hydrometeors in a significantly larger sample volume. The system measures the droplet size distribution (DSD) and number density with an added capability of discriminating between liquid water and ice hydrometeors (based on polarization-resolved side-scatter imaging). The instrument will implement automatic switching or alternating between the regular side-scatter imaging mode and sparse/large hydrometeor mode (based on the acquired data). Other adaptive sampling capabilities include variable sample volume and dynamic range (based on the measured DSD). The preferred deployment platforms are uncrewed aircraft systems (UAS) and tethered balloon/kite systems (TBS). The Phase I project achieved (or exceeded) the goals listed in the Work Plan. A CDMS-2 laboratory prototype implementing the polarization-resolved side-scatter imaging mode and direct bright-field-illumination imaging mode was designed and built. Additional capabilities included the variable illumination pulse energy and sample volume. The smallest detectable droplet diameter was improved to 3–4 μm (from the nominal 10 μm value specified for the original CDMS). Discrimination between water droplets and ice particles was experimentally demonstrated. The Phase I prototype was extensively tested and calibrated in the laboratory and also tested in the Pi Cloud Chamber at Michigan Technological University (MTU). The two intensive experimental campaigns at MTU provided unique opportunities of testing the CDMS-2 laboratory prototype under realistic warm and mixed-phase cloud conditions (stable for long periods of time), testing different sampling modes and intercomparing the CDMS-2 prototype to other co-located cloud characterization instruments. The Phase I project successfully demonstrated the feasibility of the proposed technology and identified the engineering challenges of designing a field deployable prototype instrument in Phase II. The Phase I study provides a solid basis for development, characterization and field-testing of the proposed advanced cloud probe with adaptive sampling in Phase II followed by commercialization of the technology in Phase III.

47 OTHER INSTRUMENTATION↗

An Analysis of Shallow Orographic Cumulus Clouds Observed During the CACTI Field Campaign: A SULI Internship Final Report

In 2018, the Atmospheric Radiation Measurement Aerial Facility (AAF) deployed its G-1 research aircraft to the Sierras de Córdoba mountain range in north-central Argentina to support the research of environmental factors on deep convective cycles as part of the Cloud, Aerosol, and Complex Terrain Interactions (CACTI) field campaign. The aircraft was fitted with a suite of instruments to holistically measure in situ the current state of the atmosphere. In this study, we organize the campaign’s 22 flights by environmental conditions. Data from each flight was transected by the aircraft’s position relative to cloud, allowing for in depth analysis of cloud processing on aerosol populations. My project was a case study of selected flights where warm, shallow orographic cumulus clouds were observed.

54 ENVIRONMENTAL SCIENCES↗

Exploring Saccharomycotina Yeast Ecology Through an Ecological Ontology Framework

Yeasts in the subphylum Saccharomycotina are found across the globe in disparate ecosystems. A major aim of yeast research is to understand the diversity and evolution of ecological traits, such as carbon metabolic breadth, insect association, and cactophily. This includes studying aspects of ecological traits like genetic architecture or association with other phenotypic traits. Genomic resources in the Saccharomycotina have grown rapidly. Ecological data, however, are still limited for many species, especially those only known from species descriptions where usually only a limited number of strains are studied. Moreover, ecological information is recorded in natural language format limiting high throughput computational analysis. To address these limitations, we developed an ontological framework for the analysis of yeast ecology. A total of 1,088 yeast strains were added to the Ontology of Yeast Environments (OYE) and analyzed in a machine-learning framework to connect genotype to ecology. This framework is flexible and can be extended to additional isolates, species, or environmental sequencing data. Widespread adoption of OYE would greatly aid the study of macroecology in the Saccharomycotina subphylum.

59 BASIC BIOLOGICAL SCIENCES↗

Toward equitable environmental exposure modeling through convergence of data, open, and citizen sciences: an example of air pollution exposure modeling amidst increasing wildfire smoke

Exposure modeling is critical in environmental epidemiology and human health but may face challenges (e.g., skewed data, unequal error, context-insensitive validation, and computational demands). Modeling decisions reflect the intended use of the models and the values that modelers prioritize. We aimed to provide a conceptual framework and machine learning (ML) modeling protocols that address these issues. With 500m-gridded hourly PM 2.5 and O 3 levels in Illinois before, during, and after the 2023 Canadian wildfire season as a motivating example, we conducted modeling experiments to evaluate modeling methods, guided by three domains we propose based on theories of science: 1) Data Diversity, leveraging open and citizen science data to enhance inclusivity, parsimony, and representativeness; 2) Equitable Accuracy, ensuring fairly distributed uncertainties across subpopulations; and 3) Sustainable Modeling, balancing accuracy with reducing computational demands to promote accessibility for under-resourced researchers. Here, we found that ML with publicly available data can achieve high accuracy. Depending on methods, performance may vary substantially, even with identical input data. Large but skewed data may reduce performance. Misuse of cross-validation protocols can underestimate prediction error; although we observed R 2 s of ∼98 %, the modeled estimates varied significantly, indicating the need for careful model validation. By using new modeling protocols including representativeness-considered training and validation data and a new loss function, we achieved high agreement between estimates and ground-based measurements (e.g., R 2 = ∼90 % for PM 2.5 ; ∼80 % for O 3 ), equally distributed errors across sociodemographic strata and urban–rural divides, and reduction in computation time—from several weeks or months to a few days.

Exposure assessment↗

Combining compositional data sets introduces error in covariance network reconstruction

Microbial communities are diverse biological systems that include taxa from across multiple kingdoms of life. Notably, interactions between bacteria and fungi play a significant role in determining community structure. However, these statistical associations across kingdoms are more difficult to infer than intra-kingdom associations due to the nature of the data involved using standard network inference techniques. We quantify the challenges of cross-kingdom network inference from both theoretical and practical points of view using synthetic and real-world microbiome data. We detail the theoretical issue presented by combining compositional data sets drawn from the same environment, e.g. 16S and ITS sequencing of a single set of samples, and we survey common network inference techniques for their ability to handle this error. We then test these techniques for the accuracy and usefulness of their intra- and interkingdom associations by inferring networks from a set of simulated samples for which a ground-truth set of associations is known. We show that while the two methods mitigate the error of cross-kingdom inference, there is little difference between techniques for key practical applications including identification of strong correlations and identification of possible keystone taxa (i.e. hub nodes in the network). Furthermore, we identify a signature of the error caused by transkingdom network inference and demonstrate that it appears in networks constructed using real-world environmental microbiome data.

59 BASIC BIOLOGICAL SCIENCES↗