Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “human identification”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Breaking the barrier of human-annotated training data for machine learning-aided plant research using aerial imagery

Machine learning (ML) can accelerate biological research. However, the adoption of such tools to facilitate phenotyping based on sensor data has been limited by (i) the need for a large amount of human-annotated training data for each context in which the tool is used and (ii) phenotypes varying across contexts defined in terms of genetics and environment. This is a major bottleneck because acquiring training data is generally costly and time-consuming. This study demonstrates how a ML approach can address these challenges by minimizing the amount of human supervision needed for tool building. A case study was performed to compare ML approaches that examine images collected by an uncrewed aerial vehicle to determine the presence/absence of panicles (i.e. “heading”) across thousands of field plots containing genetically diverse breeding populations of 2 Miscanthus species. Automated analysis of aerial imagery enabled the identification of heading approximately 9 times faster than in-field visual inspection by humans. Leveraging an Efficiently Supervised Generative Adversarial Network (ESGAN) learning strategy reduced the requirement for human-annotated data by 1 to 2 orders of magnitude compared to traditional, fully supervised learning approaches. The ESGAN model learned the salient features of the data set by using thousands of unlabeled images to inform the discriminative ability of a classifier so that it required minimal human-labeled training data. This method can accelerate the phenotyping of heading date as a measure of flowering time in Miscanthus across diverse contexts (e.g. in multistate trials) and opens avenues to promote the broad adoption of ML tools.

59 BASIC BIOLOGICAL SCIENCES↗

pixelvar79/ESGAN-Flowering-Detection-paper

Machine learning (ML) can accelerate biological research. However, the adoption of such tools to facilitate phenotyping based on sensor data has been limited by (i) the need for a large amount of human-annotated training data for each context in which the tool is used and (ii) phenotypes varying across contexts defined in terms of genetics and environment. This is a major bottleneck because acquiring training data is generally costly and time-consuming. This study demonstrates how a ML approach can address these challenges by minimizing the amount of human supervision needed for tool building. A case study was performed to compare ML approaches that examine images collected by an uncrewed aerial vehicle to determine the presence/absence of panicles (i.e. “heading”) across thousands of field plots containing genetically diverse breeding populations of 2 Miscanthus species. Automated analysis of aerial imagery enabled the identification of heading approximately 9 times faster than in-field visual inspection by humans. Leveraging an Efficiently Supervised Generative Adversarial Network (ESGAN) learning strategy reduced the requirement for human-annotated data by 1 to 2 orders of magnitude compared to traditional, fully supervised learning approaches. The ESGAN model learned the salient features of the data set by using thousands of unlabeled images to inform the discriminative ability of a classifier so that it required minimal human-labeled training data. This method can accelerate the phenotyping of heading date as a measure of flowering time in Miscanthus across diverse contexts (e.g. in multistate trials) and opens avenues to promote the broad adoption of ML tools.

Varela, Sebastian↗

Matrix Metalloproteinases as Candidate Antigenic Determinants for Anti‐Tumor Autoantibodies in Human Ovarian Cancer: A Post Hoc Analysis

Circulating antibodies in patients with cancer can facilitate the identification of accessible epitopes on autoantigens expressed by tumors. To identify previously unrecognized protein targets in ovarian cancer, we computationally assessed a heptapeptide consensus motif (VPELGHE, flanked by two cysteine residues yielding a cyclic nonapeptide under oxidizing conditions) previously discovered via phage display-based epitope mapping of autoantibodies in patients. Eight proteins associated with ovarian cancer encompass amino acid sequences similar to the consensus motif and were, therefore, considered as candidate native autoantigens. Among these candidate targets, however, matrix metalloproteinase 14 (MMP14) demonstrates gene expression that is both high and negatively correlated with survival in ovarian cancer patient cohorts. MMP14 protein levels are also stable in tumor versus non-tumor tissues. Moreover, the corresponding heptapeptide mimic in MMP14 occurs within an α-helical secondary structural element observed in its catalytic domain. These findings demonstrate that a subset of patient-derived autoantibodies may interact with a previously unknown antigenic epitope found in MMP14 and other MMPs, thereby providing opportunities for the development of new targeted agents.

Biochemistry & Molecular Biology↗

A derecho climatology (2004–2021) in the United States based on machine learning identification of bow echoes

Due to their persistent widespread severe winds, derechos pose significant threats to human safety and property, with impacts comparable to many tornadoes and hurricanes. Yet, automated detection of derechos remains challenging due to the absence of spatiotemporally continuous observations and the complex criteria employed to define the phenomenon. This study presents an objective derecho detection approach capable of automatically identifying derechos through both observations and model results. The approach is grounded in a physically based definition of derechos and integrates three algorithms: (1) the Python Flexible Object Tracker (PyFLEXTRKR) algorithm to track mesoscale convective systems (MCSs), (2) a semantic segmentation convolutional neural network to identify bow echoes, and (3) a comprehensive classification algorithm to detect derechos within MCS life cycles and distinguish derecho-producing from non-derecho-producing MCSs. Using this approach, we developed a novel high-resolution (4 km and hourly) observational dataset of derechos and accompanying derecho-producing MCSs over the United States east of the Rocky Mountains from 2004 to 2021. The dataset consists of two subsets based on different gust speed data sources and is analyzed to document the climatology of derechos in the United States. On average, 12–15 derechos are identified per year, aligning with previous estimations (∼6–21 events annually). The spatial distribution and seasonal variation patterns are consistent with prior studies, showing peak occurrences in the Great Plains and the Midwest during the warm season. Additionally, during the study period, derechos account for approximately 3.1 % of measured damaging gusts (≥25.93 m s−1) over the eastern United States. The dataset is publicly available at https://doi.org/10.5281/zenodo.14835362 (Li et al., 2025).

54 ENVIRONMENTAL SCIENCES↗

A randomized multiplex CRISPRi-Seq approach for the identification of critical combinations of genes

Identifying virulence-critical genes from pathogens is often limited by functional redundancy. To rapidly interrogate the contributions of combinations of genes to a biological outcome, we have developed a multiplex, randomized CRISPR interference sequencing (MuRCiS) approach. At its center is a new method for the randomized self-assembly of CRISPR arrays from synthetic oligonucleotide pairs. When paired with PacBio long-read sequencing, MuRCiS allowed for near-comprehensive interrogation of all pairwise combinations of a group of 44 Legionella pneumophila virulence genes encoding highly conserved transmembrane proteins for their role in pathogenesis. Both amoeba and human macrophages were challenged with L. pneumophila bearing the pooled CRISPR array libraries, leading to the identification of several new virulence-critical combinations of genes. lpg2888 and lpg3000 were particularly fascinating for their apparent redundant functions during L. pneumophila human macrophage infection, while lpg3000 alone was essential for L. pneumophila virulence in the amoeban host Acanthamoeba castellanii. Thus, MuRCiS provides a method for rapid genetic examination of even large groups of redundant genes, setting the stage for application of this technology to a variety of biological contexts and organisms.

79 ASTRONOMY AND ASTROPHYSICS↗

A randomized multiplex CRISPRi-Seq approach for the identification of critical combinations of genes

Identifying virulence-critical genes from pathogens is often limited by functional redundancy. To rapidly interrogate the contributions of combinations of genes to a biological outcome, we have developed a mu ltiplex, r andomized C RISPR i nterference s equencing (MuRCiS) approach. At its center is a new method for the randomized self-assembly of CRISPR arrays from synthetic oligonucleotide pairs. When paired with PacBio long-read sequencing, MuRCiS allowed for near-comprehensive interrogation of all pairwise combinations of a group of 44 Legionella pneumophila virulence genes encoding highly conserved transmembrane proteins for their role in pathogenesis. Both amoeba and human macrophages were challenged with L. pneumophila bearing the pooled CRISPR array libraries, leading to the identification of several new virulence-critical combinations of genes. lpg2888 and lpg3000 were particularly fascinating for their apparent redundant functions during L. pneumophila human macrophage infection, while lpg3000 alone was essential for L. pneumophila virulence in the amoeban host Acanthamoeba castellanii . Thus, MuRCiS provides a method for rapid genetic examination of even large groups of redundant genes, setting the stage for application of this technology to a variety of biological contexts and organisms.

59 BASIC BIOLOGICAL SCIENCES↗

A randomized multiplex CRISPRi-Seq approach for the identification of critical combinations of genes

Identifying virulence-critical genes from pathogens is often limited by functional redundancy. To rapidly interrogate the contributions of combinations of genes to a biological outcome, we have developed a mu ltiplex, r andomized C RISPR i nterference s equencing (MuRCiS) approach. At its center is a new method for the randomized self-assembly of CRISPR arrays from synthetic oligonucleotide pairs. When paired with PacBio long-read sequencing, MuRCiS allowed for near-comprehensive interrogation of all pairwise combinations of a group of 44 Legionella pneumophila virulence genes encoding highly conserved transmembrane proteins for their role in pathogenesis. Both amoeba and human macrophages were challenged with L. pneumophila bearing the pooled CRISPR array libraries, leading to the identification of several new virulence-critical combinations of genes. lpg2888 and lpg3000 were particularly fascinating for their apparent redundant functions during L. pneumophila human macrophage infection, while lpg3000 alone was essential for L. pneumophila virulence in the amoeban host Acanthamoeba castellanii . Thus, MuRCiS provides a method for rapid genetic examination of even large groups of redundant genes, setting the stage for application of this technology to a variety of biological contexts and organisms.

Ellis, Nicole A. (ORCID:0000000195928415)↗

Dara: Automated Multiple-Hypothesis Phase Identification and Refinement from Powder X-ray Diffraction

Powder X-ray diffraction (XRD) is a foundational technique for characterizing crystalline materials. However, the reliable interpretation of XRD patterns, particularly in multiphase systems, remains a manual and expertise-demanding task. As a characterization method that only provides structural information, multiple reference phases can often be fit to a single pattern, leading to potential misinterpretation when alternative solutions are overlooked. To ease humans’ efforts and address the challenge, we introduce Dara (data-driven automated Rietveld analysis), a framework designed to automate the robust identification and refinement of multiple phases from powder XRD data. Dara performs an exhaustive tree search over all plausible phase combinations within a given chemical space and validates each hypothesis using the BGMN Rietveld refinement routine. Key features include structural database filtering, automatic clustering of isostructural phases during tree expansion, and peak-matching-based scoring to identify promising phases for refinement. When ambiguity exists, Dara generates multiple hypothesis which can then be decided between by human experts or with further characterization tools. By enhancing the reliability and accuracy of phase identification, Dara enables scalable analysis of realistic complex XRD patterns and provides a foundation for integration into multimodal characterization workflows, moving toward fully self-driving materials discovery.

Biological databases↗

Genomic reconstruction of Bacillus anthracis from complex environmental samples enables high-throughput identification and lineage assignment in Pakistan

Bacillus anthracis, the causative agent of anthrax, is a highly virulent zoonotic pathogen primarily affecting domesticated and wild herbivores. Human exposure to B. anthracis is primarily through contact with infected animals or contaminated animal products. In Pakistan, where livestock vaccines are largely unavailable and infected carcasses are often disposed of improperly, the risk to humans, wildlife and livestock is significant. Currently, the diagnosis of anthrax infections and outbreak tracing necessitates the isolation and culturing of B. anthracis, a process that requires BSL-3 facilities. In this study, we show that positive identification, genome reconstruction and lineage assignment can be accomplished using bioinformatic analysis of DNA extracted directly from environmental samples that would otherwise provide the starting material for isolation and culturing. This approach does not require laboratory target enrichment as is necessary for other pathogens, due in part to the extremely high bacterial load in the bloodstream in the deceased animals. Using these methods, we greatly expand the knowledge of endemic B. anthracis in Pakistan. We provide the first reference B. anthracis genomes from Pakistan since the 1970s and identify A.Br.014 Aust94 as a minor circulating sublineage alongside the dominant A.Br.047 Vollum. Future work will focus on the limits of detection and will determine if this bioinformatic method can be expanded more broadly for B. anthracis or other pathogens to replace typical culture-based methods.

A.Br.047 Vollum↗

Long-Term Evaluation of Remedial Technology Performance in the Laboratory: Multi-year Experimental Test Plan

An understanding of the long-term effectiveness of remediation technologies is central to sustainable environmental cleanup. Long-term experiments are valuable for reducing uncertainty and predicting remediation outcomes at scales required for regulatory compliance. However, these types of tests can be costly and challenging to interpret. Therefore, contaminated sites often rely on short-term laboratory experiments that may not account for the potentially significant effects of gradual, time- dependent processes governing contaminant retention, release, and species transformation. For example, short-term lab experiments (from months to a year) conducted with sediments from the unsaturated and saturated zones at the Hanford Site play an important role in initial evaluations of remediation technologies but cannot capture the full extent of time-dependent reactions, leaving uncertainties in field-scale deployment. Here, long-term testing will be conducted on select technologies based on their performance in short-term testing. The overall objective is to directly address the challenges described above by generating and analyzing data on long-term (2–10 years) efficacy of selected remediation technologies, integrating this understanding into models, and providing critical input for remediation planning, monitoring, and 5-year review cycles for field-implemented remedies. Specific objectives include: 1. Evaluating long-term efficiency of promising remedies under site-specific conditions. 2. Generating robust parameters for modeling, reducing uncertainty in predictive simulations. 3. Advancing integrated monitoring by combining geochemical and geophysical observations. 4. Informing field-scale implementation by incrementally advancing technologies, identifying failure mechanisms early, and prioritizing robust, cost-effective remedies. Through systematic evaluation of technologies in laboratory-scale column experiments, integrated monitoring, and modeling support, this project is designed to bolster confidence in the long-term robustness of selected technologies. The ultimate outcome is the identification and deployment of more reliable, cost-effective remedies that safeguard human health and the environment while reducing the uncertainties that have historically hindered cleanup progress at Hanford Site. An experimental approach was developed and initiated for long-term testing potentially up to 10 years. The table below summarizes the experimental approach developed for testing select technologies and presented in this multi-year experimental test plan.

54 ENVIRONMENTAL SCIENCES↗

A funnel approach to enable analyses of epitope-specific human CD4 T cells specific for influenza and SARS-CoV-2

Protection against pathogens relies heavily on the adaptive immune response, whose key regulators are CD4 T cells. CD4 T cells, notable for their complex repertoire and functional potential, can most easily be dissected by identifying, quantifying, characterizing, and isolating epitope-specific cells. In the study reported here, we present a systematic and unbiased strategy that has enabled the identification of highly immunogenic peptide epitopes derived from influenza virus and SARS-CoV-2, presented by human HLA-DR proteins. Coupling the use of HLA-DR transgenic mice with infection and vaccination and highly sensitive epitope-specific cytokine ELISpot assays, we have narrowed the potential epitopes from 450 to 600 peptides to 5–15 peptides for each allele by an iterative process of elimination and selection, which we have termed a funnel approach. These epitopes have been validated in HLA-DR-typed human CD4 T cells directly ex vivo and enabled the derivation and implementation of HLA-DR peptide tetramers. Tetramer staining of human PBMCs enriched for CD4 T memory populations from healthy adult subjects, highlighted this approach as a sensitive and specific method for identifying novel epitopes, and subsequent CD4 T-cell responses to human viral infections.

CD4 T cell↗

An expedited screening platform for the discovery of anti-ageing compounds in vitro and in vivo

Background: Restraining or slowing ageing hallmarks at the cellular level have been proposed as a route to increased organismal lifespan and healthspan. Consequently, there is great interest in anti-ageing drug discovery. However, this currently requires laborious and lengthy longevity analysis. Here, we present a novel screening readout for the expedited discovery of compounds that restrain ageing of cell populations in vitro and enable extension of in vivo lifespan. Methods: Using Illumina methylation arrays, we monitored DNA methylation changes accompanying long-term passaging of adult primary human cells in culture. This enabled us to develop, test, and validate the CellPopAge Clock, an epigenetic clock with underlying algorithm, unique among existing epigenetic clocks for its design to detect anti-ageing compounds in vitro. Additionally, we measured markers of senescence and performed longevity experiments in vivo in Drosophila, to further validate our approach to discover novel anti-ageing compounds. Finally, we bench mark our epigenetic clock with other available epigenetic clocks to consolidate its usefulness and specialisation for primary cells in culture. Results: We developed a novel epigenetic clock, the CellPopAge Clock, to accurately monitor the age of a population of adult human primary cells. We find that the CellPopAge Clock can detect decelerated passage-based ageing of human primary cells treated with rapamycin or trametinib, well-established longevity drugs. We then utilise the CellPopAge Clock as a screening tool for the identification of compounds which decelerate ageing of cell populations, uncovering novel anti-ageing drugs, torin2 and dactolisib (BEZ-235). We demonstrate that delayed epigenetic ageing in human primary cells treated with anti-ageing compounds is accompanied by a reduction in senescence and ageing biomarkers. Finally, we extend our screening platform in vivo by taking advantage of a specially formulated holidic medium for increased drug bioavailability in Drosophila. We show that the novel anti-ageing drugs, torin2 and dactolisib (BEZ-235), increase longevity in vivo. Conclusions: Our method expands the scope of CpG methylation profiling to accurately and rapidly detecting anti-ageing potential of drugs using human cells in vitro, and in vivo, providing a novel accelerated discovery platform to test sought after anti-ageing compounds and geroprotectors.

60 APPLIED LIFE SCIENCES↗

Comparing Top-Down Proteoform Identification: Deconvolution, PrSM Overlap, and PTM Detection

Generating top-down tandem mass spectra (MS/MS) for complex mixtures of proteoforms has become possible through improvements in fractionation, on-line separation, dissociation, and mass analysis. The algorithms to match tandem mass spectra to sequences have undergone a parallel evolution, with both spectral alignment and peak matching being paired with diverse methods for scoring proteoform-spectral matches (PrSMs). This study assesses state-of-the-art algorithms for top-down identification through three distinct challenges. The first is identifying a large yield of PrSMs while controlling false discovery rate (FDR) in identifying thousands of proteoforms from complex cell lysates via four software workflows: ProSight Proteome Discoverer, TopPIC, Informed Proteomics, and pTop. The second is the deconvolution of data from both Thermo Orbitrap-class and Bruker maXis Q-TOF instruments to produce consistent precursor charge and mass determinations while generating fragment mass lists to optimize identification. The third attempts to detect diverse post-translational modifications (PTMs) in proteoforms from cow milk and human ovarian tissue. The data demonstrate that existing software suites produce admirable sensitivity, in some cases identifying a third of collected tandem mass spectra with FDR controlled below 2%; the overlap in these PrSMs, however, illustrates real value in searching data with multiple search engines. Differences among identification workflows seem to result from each search algorithm incorporating its own deconvolution algorithm. By transmitting deconvolution data from multiple deconvolution routes (Thermo Xtract, Bruker Auto MSn, Mascot Distiller, TopFD, and FLASHDeconv) to the downstream TopPIC search algorithm, we were able to detect common causes of deconvolution disagreement. The detection of PTMs was very inconsistent among search algorithms, with some workflows suggesting as little as 1% of PrSMs from cow’s milk were singly-phosphorylated while other workflows found that 18% of PrSMs were singly-phosphorylated. Taken together, these results make a strong argument for top-down researchers to adopt a standard practice of analyzing each MS/MS experiment with at least two different search engines.

59 BASIC BIOLOGICAL SCIENCES↗

Characterizing Families of Spectral Similarity Scores and Their Use Cases for Gas Chromatography–Mass Spectrometry Small Molecule Identification

Metabolomics provides a unique snapshot into the world of small molecules and the complex biological processes that govern the human, animal, plant, and environmental ecosystems encapsulated by the One Health modeling framework. However, this “molecular snapshot” is only as informative as the number of metabolites confidently identified within it. The spectral similarity (SS) score is traditionally used to identify compound(s) in mass spectrometry approaches to metabolomics, where spectra are matched to reference libraries of candidate spectra. Unfortunately, there is little consensus on which of the dozens of available SS metrics should be used. This lack of standard SS score creates analytic uncertainty and potentially leads to issues in reproducibility, especially as these data are integrated across other domains. In this work, we use metabolomic spectral similarity as a case study to showcase the challenges in consistency within just one piece of the One Health framework that must be addressed to enable data science approaches for One Health problems. Here, using a large cohort of datasets comprising both standard and complex datasets with expert-verified truth annotations, we evaluated the effectiveness of 66 similarity metrics to delineate between correct matches (true positives) and incorrect matches (true negatives). We additionally characterize the families of these metrics to make informed recommendations for their use. Our results indicate that specific families of metrics (the Inner Product, Correlative, and Intersection families of scores) tend to perform better than others, with no single similarity metric performing optimally for all queried spectra. This work and its findings provide an empirically-based resource for researchers to use in their selection of similarity metrics for GC-MS identification, increasing scientific reproducibility through taking steps towards standardizing identification workflows.

59 BASIC BIOLOGICAL SCIENCES↗

Methods and compositions for identification of source of microbial contamination in a sample

Herein are described 1058 different bacterial taxa that were unique to either human, grazing mammal, or bird fecal wastes. These identified taxa can serve as specific identifier taxa for these sources in environmental waters. Two field tests in marine waters demonstrate the capacity of phylogenetic microarray analysis to track multiple sources with one test.

Andersen, Gary L.↗

Machine learning-assisted elucidation of CD81–CD44 interactions in promoting cancer stemness and extracellular vesicle integrity

Tumor-initiating cells with reprogramming plasticity or stem-progenitor cell properties (stemness) are thought to be essential for cancer development and metastatic regeneration in many cancers; however, elucidation of the underlying molecular network and pathways remains demanding. Combining machine learning and experimental investigation, here we report CD81, a tetraspanin transmembrane protein known to be enriched in extracellular vesicles (EVs), as a newly identified driver of breast cancer stemness and metastasis. Using protein structure modeling and interface prediction-guided mutagenesis, we demonstrate that membrane CD81 interacts with CD44 through their extracellular regions in promoting tumor cell cluster formation and lung metastasis of triple negative breast cancer (TNBC) in human and mouse models. In-depth global and phosphoproteomic analyses of tumor cells deficient with CD81 or CD44 unveils endocytosis-related pathway alterations, leading to further identification of a quality-keeping role of CD44 and CD81 in EV secretion as well as in EV-associated stemness-promoting function. CD81 is coexpressed along with CD44 in human circulating tumor cells (CTCs) and enriched in clustered CTCs that promote cancer stemness and metastasis, supporting the clinical significance of CD81 in association with patient outcomes. Our study highlights machine learning as a powerful tool in facilitating the molecular understanding of new molecular targets in regulating stemness and metastasis of TNBC.

59 BASIC BIOLOGICAL SCIENCES↗

Top-down proteomics

Proteoforms arising from posttranslational modifications, genetic polymorphisms, and RNA splice variants, play a pivotal role as the key drivers in biology. Thus, a comprehensive understanding of proteoforms is essential for unraveling the intricacies of biological systems and bridging the gap between genotype and phenotype. By analyzing whole proteins without digestion, top-down proteomics (TDP) provides a holistic view of the proteome and presents a next-generation approach for deciphering protein function, uncovering disease mechanisms, and advancing precision medicine. This Primer embarks on a journey into the world of TDP by encapsulating its historical context, underlying principles, recent advances, and an outlook on the future of TDP. The experimental section navigates instrumentation, sample preparation, intact protein separation, tandem mass spectrometry techniques, and data collection. Results decipher raw data, visualize intact protein spectra, unravel data analysis, and explain proteoform identification, characterization, and quantitation, as well as statistical analysis. Various applications of TDP spanning the human proteoform project, biomedical, biopharmaceutical, and clinical applications are described. These are complemented by discussions on measurement reproducibility, limitations, and a forward-looking perspective outlining uncharted waters where the field can advance, and potential exciting future applications of TDP.

Roberts, David S.↗

Speaker-targeted Synthetic Speech Detection

Text-to-speech technologies are evolving quickly towards realistic-sounding human-like voices. As this technology improves, so does the opportunity for malpractice in speaker identification (SID) via spoofing, the process of impersonating a voice biometric via synthesis. More data typically equates to a more realistic voice model, which poses an issue for well-known subjects, such as politicians and celebrities, who have vast amounts of multimedia available online. Detection of synthetic speech has relied on signal processing techniques that focus on the generation of new acoustic features and train deep learning models to detect when an audio file has been manipulated through the characterization of unnatural changes or artifacts. However, these techniques do not use any information from the speaker they are evaluating. This paper proposes to incorporate information from the speaker-of-interest (SoI) into the models to avoid specific spoofing attacks for certain vulnerable people. The wealth of data for well-known people can also be used to train a speaker-specific spoofing detector with a higher level of accuracy than a speaker-independent model. The paper proposes a new xResNet-PLDA system and compares it to three different baseline systems: a state-of-the-art speaker identification system, an xResNet system trained to discriminate between bona fide and fake speech, and a speaker identification system in which the PLDA and calibration models were trained with bona fide and fake speech. We evaluated the systems in two different scenarios — a cross-validation scenario and a hold-out scenario — with three different databases. We show how the proposed system outperforms dramatically the baseline systems in each scenario and for each database. Finally, we show how using a small amount of the SoI’s speech to adapt global calibration parameters improves the performance of the system, especially in unseen conditions.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗