Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “human identification”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Ultrasensitive single-cell proteomics workflow identifies >1000 protein groups per mammalian cell

Here, we report on the combination of nanodroplet sample preparation, ultra-low-flow nanoLC, high-field asymmetric ion mobility spectrometry (FAIMS), and the latest-generation Orbitrap Eclipse Tribrid mass spectrometer for greatly improved single-cell proteome profiling. FAIMS effectively filtered out singly charged ions for more effective MS analysis of multiply charged peptides, resulting in an average of 1056 protein groups identified from single HeLa cells without MS1-level feature matching. This is 2.3 times more identifications than without FAIMS and a far greater level of proteome coverage for single mammalian cells than has been previously reported for a label-free study. Differential analysis of single microdissected motor neurons and interneurons from human spinal tissue indicated a similar level of proteome coverage, and the two subpopulations of cells were readily differentiated based on single-cell label-free quantification.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Identification and Clinical Evaluation of Potential Biomarkers for Breast Cancer Resistance Protein ( BCRP / ABCG2 )

Clinical inhibition and genetic variation of the Breast Cancer Resistance Protein (BCRP/ABCG2) efflux transporter can significantly influence drug exposure, highlighting the need for reliable BCRP functional biomarkers. This study aimed to identify and evaluate biomarkers predictive of BCRP function in humans. A comprehensive analysis of metabolomic genome‐wide association studies (mGWAS) was conducted to discover potential BCRP biomarkers, followed by evaluation inin vitrotransporter assays and a clinical drug–drug interaction (DDI) study. Across multiple mGWAS datasets, plasma concentrations of three herbicide derivatives—4‐hydroxychlorothalonil (4HC), 3‐bromo‐5‐chloro‐2,6‐dihydroxybenzoic acid (BCDBA), and 3,5‐dichloro‐2,6‐dihydroxybenzoic acid (DCDBA)—were significantly elevated (P < 5E‐8) in individuals carrying reduced functionABCG2polymorphisms. These compounds were confirmed as novel BCRP substrates via transporter uptake assays and selected for clinical evaluation alongside riboflavin, a known BCRP substrate and potential BCRP biomarker. In a DDI study with 11 healthy subjects, eltrombopag, a BCRP inhibitor, increased rosuvastatin concentrations by approximately twofold (P = 0.002). No significant changes in the plasma concentrations of organic anion transporting polypeptide 1B (OATP1B) biomarkers (CP‐I and CP‐III) or potential BCRP biomarkers (4HC, BCDBA, DCDBA, or riboflavin) were observed. Notably, two subjects were heterozygous carriers for theABCG2p.Q141K variant and exhibited significantly higher baseline concentrations of 4HC (P = 0.004) and BCDBA (P = 0.0003), consistent with reduced BCRP function. These findings suggest that 4HC and BCDBA are promising biomarkers for baseline BCRP function in specific populations, such as those harboring reduced function genetic polymorphisms, but do not appear suitable for detecting acute BCRP inhibition.

Pharmacology & Pharmacy↗

A Multiplexed Quantitative Analysis of Germline Single Amino Acid Variants by Targeted Proteomics in Nondepleted Human Plasma

Single amino acid variants (SAAVs) in protein sequences are often a direct result of single-nucleotide polymorphisms (SNPs). Certain germline SAAVs have shown biological relevance in different disease conditions but lack precise quantification in circulation, which could hinder functional investigations and progress in biomarker development. Here, we have developed a multiplexed liquid chromatography-selected reaction monitoring (LC-SRM) assay that monitors 5 wild-type and variant peptide pairs (Complement Factor B: CFB-R32Q/R32W, Clusterin: CLU-N317H, Fetuin B: FETUB-K360R, and Kininogen: KNG1-L212P) in nondepleted human plasma. The assay was optimized for imprecision, linearity, stability, and calibration assessments with CVs of under 20%. The wild-type and variant peptide pairs were characterized in a set of healthy individual plasma samples. These target identifications were also validated by SNP genotyping with more than 99% accuracy. For all protein targets, we observed significantly lower concentrations of WT species in the presence variant peptides. In CFB, the concentration of R32Q was significantly lower than its counterpart R32W variant and WT species. Furthermore, our results distinguished phenotypes of homozygosity and heterozygosity of the SAAV presence through direct concentration level characterization. These findings provide some insights into how SAAVs affect quantitative assessments of target peptides. The assay demonstrates a platform for proteogenomic analyses with potential applications in both research and clinical settings.

genetics↗

Improving Cyber Situational Understanding

Effective cybersecurity operations require the ability to analyze large amounts of information to assess security risks and formulate defensive strategies against adversaries. This has become more complex in recent years as the sprawl and interconnectivity of devices grows through implementation of virtualization, cloud computing, and Internet of Things (IoT). The amount of data and analysis required for effective cybersecurity command and control decisions far exceeds humans’ capacity to perform manually. We characterize the analysis problem as cyber situational understanding. The research presented to improve cyber situational understanding focuses on vulnerability analysis and threat intelligence. Regarding vulnerabilities, entities must analyze and plan work for between thousands and tens of thousands of software vulnerabilities annually. Entities heavily use network firewalls to limit vulnerability exposure. As a result, some of these vulnerabilities permit exposure to adversarial exploitation, whereas others are inaccessible and therefore present negligible risk of exploitation. Distinguishing between high and low risk software vulnerabilities requires a deep understanding of the vulnerability, network firewall protection, and characteristics of the targeted device. This problem is solved by extracting network service features from vulnerability data features using both machine-learning and natural language processing. Then, the network firewall topology is parsed to determine which vulnerabilities are reachable by adversaries. Ultimately, a state-based safety analysis ascertains which vulnerabilities are unsafe. A related vulnerability analysis problem occurs in cybersecurity operations when associating an entity’s hardware and software assets to public vulnerability databases. Assets often reveal hardware and software through installation artifacts and network service identification, and entities store these artifacts in inventory databases. However, software and hardware vendors apply a standard Common Platform Enumeration (CPE) naming convention when publicly reporting vulnerabilities. Associating these two datasets often requires many hours to days of manual inspection. The proposed solution automates the mapping approach of human analysts using fuzzy matching techniques, natural language processing, and, ultimately, machine learning to present a small set of recommendations for mapping the two datasets. The result significantly reduces human analysis time and reduces the occurrence of false positives in vulnerability notifications. Finally, cyber threat intelligence (CTI) requires associating cyber observable artifacts, such as IP addresses, URIs, and file hashes, with cyber threat tactics, techniques, and procedures. Unfortunately, most CTI data is compartmentalized across multiple organizations and cannot be shared due to the legal and reputational risk with cyber threat being associated with the entity. The approach to solving this problem inovlves using a distributed ledger with anonymous token spending and authentication. This allows a consortium of semi-trusted entities to share the workload of curating CTI for a threat sharing community’s cooperative benefit.

Huff, Philip↗

A framework to implement human reliability analysis during early design stages of advanced reactors

Nuclear power plants require human actions throughout their lifecycle from design, construction, operation, and decommissioning. However, for advanced reactors (e.g., Generation IV), the reliance on human intervention in safety-related actions is expected to be reduced or completely replaced by automated actions. The Probabilistic Risk Assessment (PRA) Standard for Advanced Non-LWR Nuclear Power Plants requires that the impacts of all operator actions are captured and incorporated in the risk of the modeled plant. Moreover, the Modernization of Technical Requirements for Licensing Advanced Reactors requires human reliability analysis (HRA) to be included throughout all design and PRA development stages. However, due to the lack of details during the early design stages, HRA is often postponed until the design is mature enough. Conducting HRA in later design stages, though it may be adequate in capturing pre-, at-, and post-initiators comes short of informing the design itself in the iterative design lifecycle. Hence, this paper presents a framework to include HRA during the design's early stages, pre-conceptual or conceptual. The proposed framework provides a process for the removal of operator actions that do not contribute to the risk and the identification of all key operator actions that are critical to the safety of the design. The results of this framework are then used to inform the design of those safety-related operator actions to update the design further. Then, using information from the updated design, this framework can be reapplied to investigate the impact of the design update on human reliability. The PRA model of the X-energy's pre-conceptual Xe-100 high-temperature gas-cooled pebble-bed reactor (HTGR-PB) design is used to demonstrate the approach. In the pre-conceptual Xe-100 PRA model, also called Phase 0 PRA model, human actions were considered an integral part of analyzing the plant response to different initiating events. Hence, in this paper, all possible human actions in the Xe-100 PRA model are identified, analyzed, and removed to emulate a design relying only on the available automated control systems. The preliminary results of this assessment show how safe the Xe-100 design is even without crediting any human actions. The results also list necessary sequences in which operator actions are critical to the risk profile of the design.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Burden of bacterial bloodstream infections and recent advances for diagnosis

Abstract Bloodstream infections (BSIs) and subsequent organ dysfunction (sepsis and septic shock) are conditions that rank among the top reasons for human mortality and have a great impact on healthcare systems. Their treatment mainly relies on the administration of broad-spectrum antimicrobials since the standard blood culture-based diagnostic methods remain time-consuming for the pathogen's identification. Consequently, the routine use of these antibiotics may lead to downstream antimicrobial resistance and failure in treatment outcomes. Recently, significant advances have been made in improving several methodologies for the identification of pathogens directly in whole blood especially regarding specificity and time to detection. Nevertheless, for the widespread implementation of these novel methods in healthcare facilities, further improvements are still needed concerning the sensitivity and cost-effectiveness to allow a faster and more appropriate antimicrobial therapy. This review is focused on the problem of BSIs and sepsis addressing several aspects like their origin, challenges, and causative agents. Also, it highlights current and emerging diagnostics technologies, discussing their strengths and weaknesses.

Costa, Susana P.↗

Identification of carbohydrate gene clusters obtained from in vitro fermentations as predictive biomarkers of prebiotic responses

Prebiotic fibers are non-digestible substrates that modulate the gut microbiome by promoting expansion of microbes having the genetic and physiological potential to utilize those molecules. Although several prebiotic substrates have been consistently shown to provide health benefits in human clinical trials, responder and non-responder phenotypes are often reported. These observations had led to interest in identifying, a priori, prebiotic responders and non-responders as a basis for personalized nutrition. In this study, we conducted in vitro fecal enrichments and applied shotgun metagenomics and machine learning tools to identify microbial gene signatures from adult subjects that could be used to predict prebiotic responders and non-responders. Using short chain fatty acids as a targeted response, we identified genetic features, consisting of carbohydrate active enzymes, transcription factors and sugar transporters, from metagenomic sequencing of in vitro fermentations for three prebiotic substrates: xylooligosacharides, fructooligosacharides, and inulin. A machine learning approach was then used to select substrate-specific gene signatures as predictive features. These features were found to be predictive for XOS responders with respect to SCFA production in an in vivo trial. Our results confirm the bifidogenic effect of commonly used prebiotic substrates along with inter-individual microbial responses towards these substrates. We successfully trained classifiers for the prediction of prebiotic responders towards XOS and inulin with robust accuracy (≥ AUC 0.9) and demonstrated its utility in a human feeding trial. Overall, the findings from this study highlight the practical implementation of pre-intervention targeted profiling of individual microbiomes to stratify responders and non-responders.

59 BASIC BIOLOGICAL SCIENCES↗

SNAPSHOT USA 2020: A second coordinated national camera trap survey of the United States during the COVID-19 pandemic

Managing wildlife populations in the face of global change requires regular data on the abundance and distribution of wild animals, but acquiring these over appropriate spatial scales in a sustainable way has proven challenging. Here, in this study, we present the data from Snapshot USA 2020, a second annual national mammal survey of the USA. This project involved 152 scientists setting camera traps in a standardized protocol at 1485 locations across 103 arrays in 43 states for a total of 52,710 trap-nights of survey effort. Most (58) of these arrays were also sampled during the same months (September and October) in 2019, providing a direct comparison of animal populations in 2 years that includes data from both during and before the COVID-19 pandemic. All data were managed by the eMammal system, with all species identifications checked by at least two reviewers. In total, we recorded 117,415 detections of 78 species of wild mammals, 9236 detections of at least 43 species of birds, 15,851 detections of six domestic animals and 23,825 detections of humans or their vehicles. Spatial differences across arrays explained more variation in the relative abundance than temporal variation across years for all 38 species modeled, although there are examples of significant site-level differences among years for many species. Temporal results show how species allocate their time and can be used to study species interactions, including between humans and wildlife. These data provide a snapshot of the mammal community of the USA for 2020 and will be useful for exploring the drivers of spatial and temporal changes in relative abundance and distribution, and the impacts of species interactions on daily activity patterns. There are no copyright restrictions, and please cite this paper when using these data, or a subset of these data, for publication.

54 ENVIRONMENTAL SCIENCES↗

An expanded registry of candidate cis -regulatory elements

Mammalian genomes contain millions of regulatory elements that control the complex patterns of gene expression. Previously, the ENCODE consortium mapped biochemical signals across hundreds of cell types and tissues and integrated these data to develop a registry containing 0.9 million human and 300,000 mouse candidate cis-regulatory elements (cCREs) annotated with potential functions. Here we have expanded the registry to include 2.37 million human and 967,000 mouse cCREs, leveraging new ENCODE datasets and enhanced computational methods. This expanded registry covers hundreds of unique cell and tissue types, providing a comprehensive understanding of gene regulation. Functional characterization data from assays such as STARR-seq, massively parallel reporter assay, CRISPR perturbation and transgenic mouse assays have profiled more than 90% of human cCREs, revealing complex regulatory functions. We identified thousands of novel silencer cCREs and demonstrated their dual enhancer and silencer roles in different cellular contexts. Integrating the registry with other ENCODE annotations facilitates genetic variation interpretation and trait-associated gene identification, exemplified by the identification of KLF1 as a novel causal gene for red blood cell traits. This expanded registry is a valuable resource for studying the regulatory genome and its impact on health and disease.

Moore, Jill E. [Univ. of Massachusetts, Worchester↗

Identifying Transient Candidates in the Dark Energy Survey Using Convolutional Neural Networks

The ability to discover new transient candidates via image differencing without direct human intervention is an important task in observational astronomy. For these kind of image classification problems, machine learning techniques such as Convolutional Neural Networks (CNNs) have shown remarkable success. In this work, we present the results of an automated transient candidate identification on images with CNNs for an extant data set from the Dark Energy Survey Supernova program, whose main focus was on using Type Ia supernovae for cosmology. By performing an architecture search of CNNs, we identify networks that efficiently select non-artifacts (e.g., supernovae, variable stars, AGN, etc.) from artifacts (image defects, mis-subtractions, etc.), achieving the efficiency of previous work performed with random Forests, without the need to expend any effort in feature identification. The CNNs also help us identify a subset of mislabeled images. Performing a relabeling of the images in this subset, the resulting classification with CNNs is significantly better than previous results, lowering the false positive rate by 27% at a fixed missed detection rate of 0.05.

79 ASTRONOMY AND ASTROPHYSICS↗

Differences in urban plant community compositions across an urban-rural gradient in Knoxville, TN

Urban forests, or vegetation in areas under heavy human influence, provide many ecosystem services to urban residents such as localized cooling via evapotranspiration, shade, filtering of air pollution, and the associated health benefits of natural spaces. In order to quantify the magnitude of localized cooling by trees growing in varying levels of urbanization (based on % impervious surfaces, e.g., buildings, pavement), urban forest species composition, tree size, and tree density must be characterized. As a part of Oak Ridge National Laboratory’s (ORNL) urban forest temperature study, we conducted tree censuses in five Knoxville city parks where ORNL meteorological stations are deployed. Moreover, we measured every woody plant ≥ 5 cm diameter at breast height (DBH) within a 50 m radius of each site’s meteorological station for its DBH and species identification. When possible, individuals were identified down to species. Certain genera (Quercus spp., Carya spp., Pinus spp.) were identified down to genera in interest of time. Individual and total site basal area were calculated from measured DBH data. Results show notable differences in urban plant community compositions and total woody plant basal area across sites, with more urban sites closer to downtown (West View and SEEED) having lower tree basal area than the more suburban sites (West Hills, Cumberland Estates, and Victor Ashe). We identified 54 species across all sites, with West Hills and Victor Ashe having the highest species diversity. Our results show differences in forest compositions and sizes across Knoxville, which are currently informing ORNL’s evapotranspiration estimates for each site. Data Summary: Census data for West Hills (WH), Cumberland Estates (CE), Victor Ashe (VA), West View (WV), and Socially Equal Energy Efficient Development or SEEED (SD) urban forests in Knoxville, TN, USA, including tree size based on diameter at breast height (DBH; 1.3 m), species identification (Latin and common names), and basal area per stem (BA=π×[.5*DBH]^2). Field data are summarized in this file: “Community_Composition_Data.CSV”. Site-specific data detailing each site’s coordinates, number of stems measured at DBH, average tree DBH, α-diversity (number of species present), and total site basal area (sum of individual basal areas per site) are in this file: “Site_Comparisons.CSV”.

Warren, Jeffrey [ORNL] (ORCID:0000000206804697)↗

Attention Guided Lymph Node Malignancy Prediction in Head and Neck Cancer

Accurate lymph node (LN) malignancy classification is essential for treatment target identification in head and neck cancer (HNC) radiation therapy. Given the constraints imposed by relatively small sample sizes in real-world medical applications, to classify LN malignancy status accurately, we proposed an attention-guided classification (AGC) scheme that (1) incorporates human knowledge (ie, LN contours) into model training to guide model’s “learning” direction, alleviating the critical requirement of large training samples by deep learning approaches; and (2) does not require accurate delineation of LNs in the inference stage but can highlight the discriminative region nearby the LN, which is important for malignancy determination.

62 RADIOLOGY AND NUCLEAR MEDICINE↗

Heat Transfer, Refrigeration and Heat Pumps

The Special Issue entitled “Heat Transfer, Refrigeration and heat Pumps” accepted papers covering a wide range of topics related to heat pumps, thermal energy storage, and low-Global Warming Potential (GWP) alternative refrigerants. Heat pumps play a vital role in providing space conditioning and water heating while utilizing the energy of the environment. Since heat pumps use renewable thermal energy from the environs to provide the desired utility, they contribute to the portfolio of technologies that mitigate carbon footprint. The heat pump may be considered a truly renewable technology if the electricity it uses comes entirely from a renewable source. More accurately, a heat pump is a “low carbon technology”. These perspectives make heat pumps an indispensable option for the future in order to reduce the nocuous human impact on the environment. More efficient heat pumping technologies are being developed for the residential, commercial, and industrial sectors of the economy. Another R&D thrust is heat pumps for cold climates. Hardware components, use of low-GWP refrigerants, and identification of systemic inefficiencies are active research areas. This article is a synopsis of the papers submitted to the Special Issue.

42 ENGINEERING↗

Parallel Multi-Omics in High-Risk Subjects for the Identification of Integrated Biomarker Signatures of Type 1 Diabetes

Background: Biomarkers are crucial for detecting early type-1 diabetes (T1D) and preventing significant β-cell loss before the onset of clinical symptoms. Here, we present proof-of-concept studies to demonstrate the potential for identifying integrated biomarker signature(s) of T1D using parallel multi-omics. Methods: Blood from human subjects at high risk for T1D (and healthy controls; n = 4 + 4) was subjected to parallel unlabeled proteomics, metabolomics, lipidomics, and transcriptomics. The integrated dataset was analyzed using Ingenuity Pathway Analysis (IPA) software for disturbances in the at-risk subjects compared to controls. Results: The final quadra-omics dataset contained 2292 proteins, 328 miRNAs, 75 metabolites, and 41 lipids that were detected in all samples without exception. Disease/function enrichment analyses consistently indicated increased activation, proliferation, and migration of CD4 T-lymphocytes and macrophages. Integrated molecular network predictions highlighted central involvement and activation of NF-κB, TGF-β, VEGF, arachidonic acid, and arginase, and inhibition of miRNA Let-7a-5p. IPA-predicted candidate biomarkers were used to construct a putative integrated signature containing several miRNAs and metabolite/lipid features in the at-risk subjects. Conclusions: Preliminary parallel quadra-omics provided a comprehensive picture of disturbances in high-risk T1D subjects and highlighted the potential for identifying associated integrated biomarker signatures. With further development and validation in larger cohorts, parallel multi-omics could ultimately facilitate the classification of T1D progressors from non-progressors.

60 APPLIED LIFE SCIENCES↗

CRISPR-COPIES: Web Tool

CRISPR/Cas system has emerged as a powerful genome-editing tool for metabolic engineering and human gene therapy. However, the conundrum of where to integrate heterologous genes on the chromosome using the CRISPR/Cas system remains an open question. Selecting a site for gene integration requires incorporation of complex criteria such as factors involved in CRISPR/Cas-mediated integration, genetic stability, and gene expression and therefore, usually requires strenuous characterization of sites on particular or different chromosomal locations. To address these issues, we developed CRISPR-COPIES, a COmputational Pipeline for the Identification of CRISPR/Cas-facilitated intEgration Sites. The tool applies ScaNN, a state-of-the-art model on the embedding-based nearest neighbor search for fast and accurate off-target search and can identify genome-wide intergenic sites for most bacterial and fungal genomes within minutes. This submission contains the code we developed to create a user-friendly web interface for CRISPR-COPIES (https://biofoundry.web.illinois.edu/copies/). We anticipate CRISPR-COPIES will serve as a useful tool for targeted DNA integration and aid in the characterization of synthetic biology toolkits, rapid strain construction to produce valuable biochemicals, and human gene and cell therapy.

Bioinformatics↗

Automatic identification and quantification of dense microcracks in high-performance fiber-reinforced cementitious composites through deep learning-based computer vision

Highlights: • A method is presented to detect, locate, quantify, and visualize dense microcracks in HPFRCC. • Quantification of dense microcracks is realized using deep learning method for the first time. • The presented method uses deep learning models that are trained using a realistic dataset size. • The presented method has a high computation efficiency for identifying and quantifying cracks. • The presented method provides crack width with errors up to 50 μm and a R{sup 2} value of 0.984. High-performance fiber-reinforced cementitious composites (HPFRCCs) feature high mechanical strengths, crack resistance, and durability. Under excessive loading, HPFRCCs demonstrate dense microcracks that are difficult to identify using existing methods. This study presents a computer vision method for identification, quantification, and visualization of microcracks in HPFRCCs based on deep learning. The presented method integrates multiple deep learning models and computer vision techniques in a hierarchical architecture. The crack pattern (e.g., number, width, and spacing of cracks) are automatically determined from pictures without human intervention. This study shows that the presented method achieves an accuracy of 0.992 for crack detection and an accuracy finer than 50 μm (R{sup 2} > 0.984) for quantification of crack width when deep learning models are trained using only 200 pictures of HPFRCCs and 200 pictures of conventional concrete with incorporation of data augmentation. The presented method is expected to be also applicable to other materials featuring complex cracks.

36 MATERIALS SCIENCE↗

Combining Multicolor FISH with Fluorescence Lifetime Imaging for Chromosomal Identification and Chromosomal Sub Structure Investigation

Understanding the structure of chromatin in chromosomes during normal and diseased state of cells is still one of the key challenges in structural biology. Using DAPI staining alone together with Fluorescence lifetime imaging (FLIM), the environment of chromatin in chromosomes can be explored. Fluorescence lifetime can be used to probe the environment of a fluorophore such as energy transfer, pH and viscosity. Multicolor FISH (M-FISH) is a technique that allows individual chromosome identification, classification as well as assessment of the entire genome. Here we describe a combined approach using DAPI as a DNA environment sensor together with FLIM and M-FISH to understand the nanometer structure of all 46 chromosomes in the nucleus covering the entire human genome at the single cell level. Upon DAPI binding to DNA minor groove followed by fluorescence lifetime measurement and imaging by multiphoton excitation, structural differences in the chromosomes can be studied and observed. This manuscript provides a blow by blow account of the protocol required to perform M-FISH-FLIM of whole chromosomes.

59 BASIC BIOLOGICAL SCIENCES↗

Defect detection in atomic-resolution images via unsupervised learning with translational invariance

Abstract Crystallographic defects can now be routinely imaged at atomic resolution with aberration-corrected scanning transmission electron microscopy (STEM) at high speed, with the potential for vast volumes of data to be acquired in relatively short times or through autonomous experiments that can continue over very long periods. Automatic detection and classification of defects in the STEM images are needed in order to handle the data in an efficient way. However, like many other tasks related to object detection and identification in artificial intelligence, it is challenging to detect and identify defects from STEM images. Furthermore, it is difficult to deal with crystal structures that have many atoms and low symmetries. Previous methods used for defect detection and classification were based on supervised learning, which requires human-labeled data. In this work, we develop an approach for defect detection with unsupervised machine learning based on a one-class support vector machine (OCSVM). We introduce two schemes of image segmentation and data preprocessing, both of which involve taking the Patterson function of each segment as inputs. We demonstrate that this method can be applied to various defects, such as point and line defects in 2D materials and twin boundaries in 3D nanocrystals.

36 MATERIALS SCIENCE↗