Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Improv”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Single crystal purification reduces trace impurities in halide perovskite precursors, alters perovskite thin film performance, and improves phase stability

Impurities present in commercially available halide perovskite precursors are known to affect photovoltaic performance. Here, we employ bulk single crystal growth of FAPbI 3 using solvent orthogonality induced crystallization (SONIC) to remove a broad set of extrinsic impurities from commercially available halide perovskite precursors, as verified by detailed chemical analysis. Following SONIC purification, FAPbI 3 films made from PbI 2 of originally low purity (99%) and high purity (99.99%) show improved phase purity and stability under light and heat relative to films made from raw precursors or precursors purified via retrograde powder crystallization (RPC) in 2-methoxyethanol, a method commonly utilized in recent reports of the highest-efficiency perovskite solar cells. In conclusion, single-crystal purification of precursors improves film stability under operational stressors, and the large enhancements in material purity provide a cleaner slate for improved isolation of compositional and additive effects on perovskite phase stability.

14 SOLAR ENERGY↗

High-density CRISPRi screens reveal diverse routes to improved acclimation in cyanobacteria

Cyanobacteria are the oldest form of photosynthetic life on Earth and contribute to primary production in nearly every habitat, from permafrost to hot springs. Despite longstanding interest in the acclimation of these microbes, it remains poorly understood and challenging to rewire. Here, this study uses a high-density, genome-wide CRISPR interference screen to examine the influence of gene-specific transcriptional variation on the growth of Synechococcus sp. PCC 7002 under environmental extremes. Surprisingly, many partial knockdowns enhanced fitness under cold monochromatic conditions. Transcriptional repression of genes for core subunits of the NDH-1 complex, which are important for photosynthesis and carbon uptake, improved growth rates under both red and blue light but at distinct, color-specific optima. Most genes with fitness-improving knockdowns were distinct to each light color, and dual-target transcriptional repression produced nonadditive effects. Findings reveal diverse routes to improved acclimation in cyanobacteria (e.g., attenuation of genes involved in CO 2 uptake, light harvesting, translation, and purine metabolism) and provide an approach for using gradients in sgRNA activity to pinpoint biochemically influential transcriptional changes in cells.

59 BASIC BIOLOGICAL SCIENCES↗

Coupled machine learning–ecosystem ensemble models substantially improve predictions of nitrous oxide (N 2 O) fluxes from US croplands

Nitrous oxide (N 2 O) is a potent and persistent greenhouse gas, with rising atmospheric concentrations driven in part by inefficient use of synthetic nitrogen (N) fertilizers in agriculture. Predicting soil N 2 O emissions is challenging due to high spatial and temporal variability arising from complex soil biogeochemical processes. Process-based ecosystem models and standalone machine learning (ML) approaches without extensive site-specific calibration often miss high-emission episodes. Here, we show how an Ensemble Modeling System (EMS) based on outputs from an ensemble of ecosystem models coupled to an ensemble of ML models can improve predictions and understanding of N 2 O fluxes from US cropland. Trained and validated on ~12,000 N 2 O chamber measurements at 17 US Midwest sites (six crops, 35 management practices), the EMS accurately predicted daily fluxes of N 2 O at both training (R 2 = 0.84, RMSE = 16.4 g N ha −1 d −1 ) and held-out testing sites (R 2 = 0.84, RMSE = 6.2 g N ha −1 d −1 ). Analyses identified six dominant N 2 O drivers: soil organic carbon (SOC), NH 4 + , NO 3 - , water-filled pore space, temperature, and aboveground biomass production. Wet, warm soils produced large N 2 O peaks only with sufficient SOC and mineral N; in low-SOC soils, fluxes remained low. Incorporating these drivers into process-based models might significantly improve their predictive capacity. The EMS demonstrates a strong potential to predict N 2 O fluxes at unseen sites, enabling more reliable regional inventories, improved gap-filling where measurements are sparse, and enhanced understanding of mechanisms to advance targeted mitigation strategies in food, feed, and bioenergy crops.

AI↗

Improving luminescence response in ZnGeN 2 /GaN superlattices: defect reduction through composition control

Abstract Color-mixed (cm) light-emitting diodes (LEDs) are theoretically the most efficient white light emitters, projected to improve white light luminous efficacy by 34% compared to incumbent phosphor converted LEDs. Since white light technology is pervasive and essential, small improvements in LED technology can result in energy savings. However, cm-LEDs are not yet realized due to poor efficacy in green and amber emitting materials, a spectral region colloquially referred to as the Green Gap. ZnGeN 2 is nearly isostructural and closely lattice-matched to GaN and can be heteroepitaxially integrated with existing GaN devices; ZnGeN 2 /GaN hybrid structures are theorized to emit green (~530 nn) light with a spontaneous emission rate 4.6–4.9 times higher than traditional InGaN LEDs when incorporated into III-N LED structures. In this report we demonstrate the molecular beam epitaxy (MBE) growth of GaN and ZnGeN 2 superlattices, an important step towards realizing multiple quantum well structures required for efficient LEDs. Elemental analysis, including atom probe tomography, shows that Ga and Ge are observed in both ZnGeN 2 and GaN layers, degrading the structural uniformity. The lack of elemental abruptness also leads to increased defect luminescence and reabsorption of band edge luminescence. The source of unintentional Ga distributed throughout the ZnGeN 2 layers was identified as excess flux escaping from around the closed MBE shutter. The source of unintentional Ge, which tended to incorporate as a single delta-doped layer in GaN, was identified as Ge riding along the cyclical metal-rich Ga adlayer used for high quality GaN, incorporating during subsequent nitrogen-rich growth step. Modifying the growth strategy results in improved structural quality, elemental abruptness, and luminescence response. This realization of structurally and elementally abrupt interfaces demonstrates the potential of heteroepitaxially integrated binary and ternary nitrides for energy-relevant devices.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Improving constraints on inflation with CMB delensing

Abstract The delensing of cosmic microwave background (CMB) maps will be increasingly valuable for extracting as much information as possible from future CMB surveys. Delensing provides many general benefits, including sharpening of the acoustic peaks, more accurate recovery of the damping tail, and reduction of lensing-inducedB-mode power. In this paper we present several applications of delensing focused on testing theories of early-universe inflation with observations of the CMB. We find that delensing the CMB results in improved parameter constraints for reconstructing the spectrum of primordial curvature fluctuations, probing oscillatory features in the primordial curvature spectrum, measuring the spatial curvature of the universe, and constraining several different models of isocurvature perturbations. In some cases we find that delensing can recover almost all of the constraining power contained in unlensed spectra, and it will be a particularly valuable analysis technique to achieve further improvements in constraints for model parameters whose measurements are not expected to improve significantly when utilizing only lensed CMB maps from next-generation CMB surveys. We also quantify the prospects of testing the single-field inflation tensor consistency condition using delensed CMB data; we find it to be out of reach of current and proposed experimental technology and advocate for alternative detection methods.

Astronomy & Astrophysics↗

Using active learning to improve quasar identification for the DESI spectra processing pipeline

The Dark Energy Spectroscopic Instrument (DESI) survey uses an automatic spectral classification pipeline to classify spectra. QuasarNET is a convolutional neural network used as part of this pipeline originally trained using data from the Baryon Oscillation Spectroscopic Survey (BOSS). In this paper we implement an active learning algorithm to optimally select spectra to use for training a new version of the QuasarNET weights file using only DESI data, with the goal of improving classification accuracy. This active learning algorithm includes a novel outlier rejection step using a Self-Organizing Map to ensure we label spectra representative of the larger quasar sample observed in DESI. We perform two iterations of the active learning pipeline, assembling a final dataset of 5600 labeled spectra, a small subset of the approximately 1.3 million quasar targets in DESI's Data Release 1. When splitting the spectra into training and validation subsets we achieve similar performance to the previously trained weights file in completeness and purity calculated on the validation dataset but do so with less than one tenth of the amount of training data. The new weights also more consistently classify objects in the same way when used on unlabeled data compared to the old weights file. In the process of improving QuasarNET's classification accuracy we discovered a systemic error in QuasarNET's redshift estimation and used our findings to improve our understanding of QuasarNET's redshifts.

Machine learning↗

Improved modeling of in-ice particle showers for IceCube event reconstruction

The IceCube Neutrino Observatory relies on an array of photomultiplier tubes to detect Cherenkov light produced by charged particles in the South Pole ice. IceCube data analyses depend on an in-depth characterization of the glacial ice, and on novel approaches in event reconstruction that utilize fast approximations of photoelectron yields. Here, a more accurate model is derived for event reconstruction that better captures our current knowledge of ice optical properties. When evaluated on a Monte Carlo simulation set, the median angular resolution for in-ice particle showers improves by over a factor of three compared to a reconstruction based on a simplified model of the ice. The most substantial improvement is obtained when including effects of birefringence due to the polycrystalline structure of the ice. When evaluated on data classified as particle showers in the high-energy starting events sample, a significantly improved description of the events is observed.

47 OTHER INSTRUMENTATION↗

Advancements and opportunities to improve bottom–up estimates of global wetland methane emissions

Wetlands are the single largest natural source of atmospheric methane (CH 4 ), contributing approximately 30% of total surface CH 4 emissions, and they have been identified as the largest source of uncertainty in the global CH 4 budget based on the most recent Global Carbon Project CH 4 report. High uncertainties in the bottom–up estimates of wetland CH 4 emissions pose significant challenges for accurately understanding their spatiotemporal variations, and for the scientific community to monitor wetland CH 4 emissions from space. In fact, there are large disagreements between bottom–up estimates versus top–down estimates inferred from inversion of atmospheric CH 4 concentrations. To address these critical gaps, we review recent development, validation, and applications of bottom–up estimates of global wetland CH 4 emissions, as well as how they are used in top–down inversions. These bottom–up estimates, using (1) empirical biogeochemical modeling (e.g. WetCHARTs: 125–208 TgCH 4 yr -1 ); (2) process-based biogeochemical modeling (e.g. WETCHIMP: 190 ± 39 TgCH 4 yr -1 ); and (3) data-driven machine learning approach (e.g. UpCH4: 146 ± 43 TgCH 4 yr -1 ). Bottom–up estimates are subject to significant uncertainties (~80 Tg CH 4 yr -1 ), and the ranges of different estimates do not overlap, further amplifying the overall uncertainty when combining multiple data products. These substantial uncertainties highlight gaps in our understanding of wetland CH 4 biogeochemistry and wetland inundation dynamics. Major tropical and arctic wetland complexes are regional hotspots of CH 4 emissions. However, the scarcity of satellite data over the tropics and northern high latitudes offer limited information for top–down inversions to improve bottom–up estimates. Recent advances in surface measurements of CH 4 fluxes (e.g. FLUXNET-CH 4 ) across a wide range of ecosystems including bogs, fens, marshes, and forest swamps provide an unprecedented opportunity to improve existing bottom–up estimates of wetland CH 4 estimates. We suggest that continuous long-term surface measurements at representative wetlands, high fidelity wetland mapping, combined with an appropriate modeling framework, will be needed to significantly improve global estimates of wetland CH 4 emissions. There is also a pressing unmet need for fine-resolution and high-precision satellite CH 4 observations directed at wetlands.

54 ENVIRONMENTAL SCIENCES↗

Improving seasonal precipitation forecasts in the Western United States through statistical downscaling

Abstract Seasonal precipitation forecasts in the western United States are critical resources for water resource management, especially during winter. While current seasonal forecasting systems provide monthly precipitation forecasts operationally, their coarse resolution limits their effectiveness in capturing the localized precipitation patterns and snowpack conditions essential for water resource managers in the mountainous regions. Here, analog statistical downscaling is demonstrated as an effective approach to enhance the spatial resolution of operational seasonal forecasts provided by the North American Multi-Model Ensemble. Downscaling was performed by building an analog ‘library’, in which corresponding model forecasts and observed values during the training period were stored. In the testing period, unseen model forecasts referenced the closest historical forecast from the analog library and applied the corresponding observational value for each point. This analysis indicates that downscaled products can capture localized features more accurately than the original coarse resolution forecasts, reducing forecast error across the western United States. Moreover, downscaling individual ensemble members—rather than downscaling the ensemble mean—further reduces forecasting error for their multi-model ensemble mean products. The greatest error reductions in the downscaled product, measured by root mean squared error (RMSE), were observed at low to mid-elevations (500–2000 meters), with 50%–70% improvement relative to the original forecasts. In the higher elevations (2000 meters and above), changes in RMSE relative to the original forecast were limited to 10%–30% improvements. The improvement is more substantial for forecast systems with 10 ensemble members compared to that with 4 members, but this relationship does not hold for the system with 24 ensemble members. These findings show that analog statistical downscaling can effectively address the spatial limitations of seasonal precipitation forecasts with minimal computational cost, providing a valuable framework for enhancing coarse resolution forecasting products while providing insights into the timing of ensemble mean calculations during the downscaling process.

Vernon, B. (ORCID:0009000891670689)↗

Improving neutrino energy estimation of charged-current interaction events with recurrent neural networks in MicroBooNE

We present a deep learning-based method for estimating the neutrino energy of charged-current neutrino-argon interactions. We employ a recurrent neural network (RNN) architecture for neutrino energy estimation in the MicroBooNE experiment, utilizing liquid argon time projection chamber (LArTPC) detector technology. Traditional energy estimation approaches in LArTPCs, which largely rely on reconstructing and summing visible energies, often experience sizable biases and resolution smearing because of the complex nature of neutrino interactions and the detector response. The estimation of neutrino energy can be improved after considering the kinematics information of reconstructed final-state particles. Utilizing kinematic information of reconstructed particles, the deep learning-based approach shows improved resolution and reduced bias for the muon neutrino Monte Carlo simulation sample compared to the traditional approach. In order to address the common concern about the effectiveness of this method on experimental data, the RNN-based energy estimator is further examined and validated with dedicated data-simulation consistency tests using MicroBooNE data. We also assess its potential impact on a neutrino oscillation study after accounting for all statistical and systematic uncertainties and show that it enhances physics sensitivity. This method has good potential to improve the performance of other physics analyses. Published by the American Physical Society 2024

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Deep-learning methods for contrast enhancement and artifact reduction in cryo-electron tomography: a systematic analysis of the state of the art and proposed improvements

Cryo-electron tomography (cryo-ET) has emerged as the preferred technique for visualizing the organization of macromolecular complexes in situ and resolving their structures at subnanometre resolution [Tegunov et al. (2021)View full citation, Nat. Methods, 18, 186–193]. Despite improvements in data quality as a result of advances in detector technology, microscope stability and stage precision, the analysis and interpretation of tomograms remains challenging due to a low signal-to-noise ratio and reconstruction artifacts stemming from experimental constraints in specimen tilt during data collection resulting in a missing wedge in the Fourier space. Recently, self-supervised deep-learning methods have been proposed for contrast enhancement and reduction of resolution anisotropy in reconstructed tomograms. Here, we evaluate several state-of-the-art deep-learning methods which aim to improve the interpretability of cryo-ET reconstructions, with a focus on their performance on downstream tasks of template matching, sub­tomogram averaging and segmentation. We propose new training architectures and a loss function based on Fourier shell correlation that show improved performance over the standard U-Net with L1/L2 losses. We demonstrate our analysis on four diverse experimental datasets: purified 80S ribosomes, in situ Chlamydomonas reinhardtii, immature HIV-1 virus-like particles and INS-1E cells.

contrast enhancement↗

Improved VOC in RbF-Treated Cu(In,Ga)Se2 Solar Cells via Passivation of Recombination Centers

Cu(In,Ga)Se 2 (CIGS) solar cells have benefited in recent years from the addition of heavy alkali elements, such as Rb, which increase the solar cell open-circuit voltage ( V OC ). To investigate the source of this improvement, here, we compare samples with and without Rb to perform a quantitative comparison of electronic defects and minority carrier lifetime. Deep-level transient and optical spectroscopy measurements were performed on two sets of rubidium fluoride (RbF)-treated and untreated CIGS, and three distinct traps were identified regardless of RbF treatment. The RbF treatment was found to reduce the concentration of the H2 trap, which was previously found to act as a recombination center and is located preferentially at CIGS grain boundaries. Time-resolved photoluminescence measurements showed an increase in effective lifetime after RbF and nearly all lifetime improvement resulted from reductions in bulk recombination. The observed V OC improvement is well correlated with increased minority carrier lifetime and acceptor concentration, which led to increases and decreases in electron and hole quasi-Fermi levels, respectively.

Cu(In Ga)Se2 (CIGS)↗

Improved PV System Control Strategies to Reduce Power Management Costs in Nanogrids

An improved PV system control method is proposed to reduce nanogrid operation costs in this paper. A model including it various components such as photovoltaic (PV) systems, energy storage systems (ESSs), gateways, and household loads is considered with the constraints of the ESS, PV irradiance from real data, and household load. We designed an optimal economic dispatch strategy with an improved PV system control method. Combining the proposed optimal economic dispatch and PV system control strategies, it can improve the control performance for both transient and steady-state responses thereby enabling the maximum power to be extracted from the PV. Consequently, the PV power is maximized, which allows the ESS to use less power and sell the surplus to external power sources, which means the proposed method decreases the nanogrid operation costs. Furthermore, this performance is verified via nanogrid simulations and PV experimental kit.

PV system control↗

Improving precision and accuracy of genetic mapping with genotyping‐by‐sequencing data in outcrossing species

Abstract Genotyping‐by‐sequencing (GBS) is a widely used strategy for obtaining large numbers of genetic markers in model and non‐model organisms. In crop plants, GBS‐derived marker datasets are frequently used to perform quantitative trait locus (QTL) mapping. In some plant species, however, high heterozygosity and complex genome structure mean that researchers must use care in handling GBS data to conduct QTL mapping most effectively. Such outbred crops include most of the perennial grass and tree species used for bioenergy. To identify strategies for increasing accuracy and precision of QTL mapping using GBS data in outbred crops, we conducted an empirical study of SNP‐calling and genetic map‐building pipeline parameters in a Miscanthus sinensis population, and a complementary simulation study to estimate the relationship between genome‐wide error rate, read depth, and marker number. The bioenergy grass Miscanthus is an obligate outcrossing species with a recent (diploidized) whole‐genome duplication. For the study of empirical M. sinensis data, we compared two SNP‐calling methods (one non‐reference‐based and one reference‐based), a series of depth filters (12×, 20×, 30×, and 40×) and two map‐construction methods (i.e., marker ordering: linkage‐only and order‐corrected based on a reference genome). We found that correcting the order of markers on a linkage map by using a high‐quality reference genome improved QTL precision (shorter confidence intervals). For typical GBS datasets of between 1000 and 5000 markers to build a genetic map for biparental populations, a depth filter set at 30× to 40× applied to outbred populations provided a genome‐wide genotype‐calling error rate of less than 1%, improved accuracy of QTL point estimates and minimized type I errors for identifying QTL. Based on these results, we recommend using a reference genome to correct the marker order of genetic maps and a robust genotype depth filter to improve QTL mapping for outbred crops.

59 BASIC BIOLOGICAL SCIENCES↗

Turbo‐charging crop improvement: harnessing multiplex editing for polygenic trait engineering and beyond

Multiplex CRISPR editing has emerged as a transformative platform for plant genome engineering, enabling the simultaneous targeting of multiple genes, regulatory elements, or chromosomal regions. This approach is effective for dissecting gene family functions, addressing genetic redundancy, engineering polygenic traits, and accelerating trait stacking and de novo domestication. Its applications now extend beyond standard gene knockouts to include epigenetic and transcriptional regulation, chromosomal engineering, and transgene‐free editing. These capabilities are advancing crop improvement not only in annual species but also in more complex systems such as polyploids, undomesticated wild relatives, and species with long generation times. At the same time, multiplex editing presents technical challenges, including complex construct design and the need for robust, scalable mutation detection. We discuss current toolkits and recent innovations in vector architecture, such as promoter and scaffold engineering, that streamline workflows and enhance editing efficiency. High‐throughput sequencing technologies, including long‐read platforms, are improving the resolution of complex editing outcomes such as structural rearrangements—often missed by standard genotyping—when targeting repetitive or tandemly spaced loci. To fully realize the potential of multiplex genome engineering, there is growing demand for user‐friendly, synthetic biology‐compatible, and scalable computational workflows for gRNA design, construct assembly, and mutation analysis. Experimentally validated inducible or tissue‐specific promoters are also highly desirable for achieving spatiotemporal control. As these tools continue to evolve, multiplex CRISPR editing is poised to become a foundational technology of next‐generation crop improvement to address challenges in agriculture, sustainability, and climate resilience.

59 BASIC BIOLOGICAL SCIENCES↗

An Alternative Ensemble Streamflow Prediction Approach Using Improved Subseasonal Precipitation Forecasts from the North America Multi-Model Ensemble Phase II

In this article, streamflow forecasting at a subseasonal time scale (10–30 days into the future) is important for various human activities. The ensemble streamflow prediction (ESP) is a widely applied technique for subseasonal streamflow forecasting. However, ESP’s reliance on the randomly resampled historical precipitation limits its predictive capability. Available dynamical subseasonal precipitation forecasts provide an alternative to the randomly resampled precipitation in ESP. Prior studies found the predictive performance of raw subseasonal precipitation forecast is limited in many regions such as the central south of the United States, which raises questions about its effectiveness in assisting streamflow forecasting. To further assess the hydrologic applicability of dynamical subseasonal precipitation forecasts, we test the subseasonal precipitation forecast from North America Multi-Model Ensemble Phase II (NMME-2) at four watersheds in the central south region of the United States. The subseasonal precipitation forecasts are postprocessed with bias correction and spatial disaggregation (BCSD) to correct bias and improve spatial resolution before replacing the randomly resampled precipitation in ESP for streamflow predictions. The performance of the resulting streamflow predictions is benchmarked with ESP. Evaluation is conducted using Kling–Gupta Efficiency (KGE), continuous ranked probability score (CRPS), probability of detection (POD), false alarm ratios (FARs), as well as reliability diagrams. Our results suggest that BCSD-corrected subseasonal precipitation forecasts lead to overall improved streamflow predictions due to added skills in winter and spring. Our results also suggest that BCSD-corrected subseasonal precipitation forecasts lead to improved predictions on the occurrence of high-percentile streamflow values above 75%. Overall, BCSD-corrected subseasonal precipitation has shown promising performance, highlighting its potential broader applications for river and flood forecasting.

54 ENVIRONMENTAL SCIENCES↗

An improved dataset for predicting mammal infecting viruses from genetic sequence information

There have been several attempts to develop machine learning (ML) models to identify human infecting viruses from their genomic sequences, with varying degrees of success. Direct comparison between models is problematic, because these models are typically trained and evaluated on different datasets with alternative data splitting schemes, features, and model performance metrics. In this paper we present a standardized dataset of mammal infecting and non-infecting viral pathogens, refined from the previous work of Mollentze et al. to include the latest literature evidence, roughly doubling the number of curated host-virus records available to the community, and new host target labels, primate and mammal. The new host labels were included for several reasons, including previous reports that classification performance is better at broader taxonomic ranks and the idea that there may be more data for primate infection that might serve as a suitable proxy for zoonotic potential and avoidance of false positives for human infection due to absence of evidence. On this dataset, we report the performance of eight machine learning models for predicting mammal-infecting viruses from their genomic sequences. We find that randomly assigning cases in our improved dataset to training/testing sets, when compared to the original assignments into training/testing in Mollentze et al., increases the overall average ROC AUC of prediction of human infection from 0.663 ± 0.070 to 0.784 ± 0.013, consistent with the reduction in phylogenetic distance between train and test sets (relative entropy change from 3.00 to 0.08). The broadest host category of mammal infection can be predicted most reliably at 0.850 ± 0.020. We share our improved dataset and code to enable standardized comparisons of machine learning methods to predict human host infections. Overall, we have presented preliminary evidence that classification of virus host infection is more tractable at higher taxonomic ranks, that unsurprisingly reducing the phylogenetic distance between training and test sets can improve predictive performance, that peptide kmer features appear to be harmful to out of sample model performance, and we are left with the question of whether models for virus host prediction can reasonably be expected to perform well in out of sample scenarios given the likelihood that viruses do not share a common ancestor. Consistent with this concern, when the data is resampled such that there is no overlap between viral families in training and test sets (relative entropy > 24), models perform no better than random chance at prediction of human infection regardless of whether kmers are included (ROC AUC 0.50 ± 0.08) or not (ROC AUC 0.50 ± 0.04).

59 BASIC BIOLOGICAL SCIENCES↗

Multi-contrast machine learning improves schistosomiasis diagnostic performance

Schistosomiasis currently affects over 250 million people and remains a public health burden despite ongoing global control efforts. Conventional microscopy is a practical tool for diagnosis and screening ofSchistosoma haematobium, but identification of eggs requires a skilled microscopist. Here we present a machine learning (ML)-based strategy for automated detection ofS. haematobiumthat combines two imaging contrasts, brightfield (BF) and darkfield (DF), to improve diagnostic performance. We collected BF and DF images of urine samples, many of them containingS. haematobiumeggs, during two different field studies in Côte d’Ivoire using a mobile phone-based microscope, the SchistoScope. We then trained separate egg-detection ML models and compared the patient-level performance of BF and DF models alone to combinations of BF and DF models, using annotations from trained microscopists as the gold standard. We found that models trained on DF images, and almost all BF and DF combinations, performed significantly better than models trained on BF images only. When models were trained on images from the first field study (n = 349 patients, 748 images of each contrast), patient-level classification performance on patient images from the second study (n = 375 patients, 752 images of each contrast) met the WHO Diagnostic Target Product Profile (TPP) sensitivity and specificity for the monitoring and evaluation use case (sensitivity for all models and combinations was >75% when evaluated at a confidence score threshold that resulted in specificity >96.5%). When we used images from both field studies for the training set, performance of the models was improved. Overall, this work shows that the use of DF and BF increases the performance of ML models on images from devices with low-cost optics, while retaining the portability, power, and time-to-results of the WHO’s diagnostic TPP. DF requires no additional sample preparation and does not increase the complexity of the imaging system. It thus offers a practical means to improve performance of automated diagnostics forS. haematobiumas well as other microscopy-based diagnostics.

Infectious Diseases↗