Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “large data sets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Investigating the Role of Copper in Arsenic Doped Cd(Se,Te) Photovoltaics

The open circuit voltage (VOC) deficit in Cd(Se,Te)-based photovoltaics remains a critical obstacle for pushing the technology closer to theoretical performance limits. Arsenic doping has become a dominant and promising route to achieve the higher p-type carrier concentrations necessary for higher VOC, but challenges associated with this alternate defect chemistry and higher doping density have hindered progress. Here we show that while arsenic doping enables high carrier concentrations (>1016 cm-3), co-doping with copper can provide a boost to VOC without a significant change to carrier concentration. A large data set is initially used to explore current-voltage and capacitance-voltage trends associated with arsenic doped devices with and without copper. A smaller subset is then used to probe these trends using a wide variety of characterization techniques. Copper is found to facilitate reduced interface recombination and potentially improved bulk absorber characteristics, though the mechanisms for these improvements are not yet clear. Despite the improved performance of co-doped devices, VOC is still far below its potential especially for highly doped devices. Low emitter doping in conjunction with high absorber doping seems to be a plausible cause for this significant deficit, though other device properties may exacerbate this problem.

CdSeTe↗

Deep Learning-Enabled MS/MS Spectrum Prediction Facilitates Automated Identification Of Novel Psychoactive Substances

The market for illicit drugs has been reshaped by the emergence of more than 1100 new psychoactive substances (NPS) over the past decade, posing a major challenge to the forensic and toxicological laboratories tasked with detecting and identifying them. Tandem mass spectrometry (MS/MS) is the primary method used to screen for NPS within seized materials or biological samples. The most contemporary workflows necessitate labor-intensive and expensive MS/MS reference standards, which may not be available for recently emerged NPS on the illicit market. Here, we present NPS-MS, a deep learning method capable of accurately predicting the MS/MS spectra of known and hypothesized NPS from their chemical structures alone. NPS-MS is trained by transfer learning from a generic MS/MS prediction model on a large data set of MS/MS spectra. We show that this approach enables a more accurate identification of NPS from experimentally acquired MS/MS spectra than any existing method. We demonstrate the application of NPS-MS to identify a novel derivative of phencyclidine (PCP) within an unknown powder seized in Denmark without the use of any reference standards. We anticipate that NPS-MS will allow forensic laboratories to identify more rapidly both known and newly emerging NPS. NPS-MS is available as a web server at https://nps-ms.ca/, which provides MS/MS spectra prediction capabilities for given NPS compounds. Additionally, it offers MS/MS spectra identification against a vast database comprising approximately 8.7 million predicted NPS compounds from DarkNPS and 24.5 million predicted ESI-QToF-MS/MS spectra for these compounds.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Predicting Small Molecule Transfer Free Energies by Combining Molecular Dynamics Simulations and Deep Learning

Accurately predicting small molecule partitioning and hydrophobicity is critical in the drug discovery process. There are many heterogeneous chemical environments within a cell and entire human body. For example, drugs must be able to cross the hydrophobic cellular membrane to reach their intracellular targets, and hydrophobicity is an important driving force for drug–protein binding. Atomistic molecular dynamics (MD) simulations are routinely used to calculate free energies of small molecules binding to proteins, crossing lipid membranes, and solvation but are computationally expensive. Machine learning (ML) and empirical methods are also used throughout drug discovery but rely on experimental data, limiting the domain of applicability. We present atomistic MD simulations calculating 15,000 small molecule free energies of transfer from water to cyclohexane. This large data set is used to train ML models that predict the free energies of transfer. We show that a spatial graph neural network model achieves the highest accuracy, followed closely by a 3D-convolutional neural network, and shallow learning based on the chemical fingerprint is significantly less accurate. A mean absolute error of ~4 kJ/mol compared to the MD calculations was achieved for our best ML model. We also show that including data from the MD simulation improves the predictions, tests the transferability of each model to a diverse set of molecules, and show multitask learning improves the predictions. This work provides insight into the hydrophobicity of small molecules and ML cheminformatics modeling, and our data set will be useful for designing and testing future ML cheminformatics methods.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Large Scale Study of Ligand–Protein Relative Binding Free Energy Calculations: Actionable Predictions from Statistically Robust Protocols

The accurate and reliable prediction of protein–ligand binding affinities can play a central role in the drug discovery process as well as in personalized medicine. Of considerable importance during lead optimization are the alchemical free energy methods that furnish an estimation of relative binding free energies (RBFE) of similar molecules. Recent advances in these methods have increased their speed, accuracy, and precision. This is evident from the increasing number of retrospective as well as prospective studies employing them. However, such methods still have limited applicability in real-world scenarios due to a number of important yet unresolved issues. Here, we report the findings from a large data set comprising over 500 ligand transformations spanning over 300 ligands binding to a diverse set of 14 different protein targets which furnish statistically robust results on the accuracy, precision, and reproducibility of RBFE calculations. We use ensemble-based methods which are the only way to provide reliable uncertainty quantification given that the underlying molecular dynamics is chaotic. These are implemented using TIES (Thermodynamic Integration with Enhanced Sampling). Results achieve chemical accuracy in all cases. Ensemble simulations also furnish information on the statistical distributions of the free energy calculations which exhibit non-normal behavior. We find that the “enhanced sampling” method known as replica exchange with solute tempering degrades RBFE predictions. We also report definitively on numerous associated alchemical factors including the choice of ligand charge method, flexibility in ligand structure, and the size of the alchemical region including the number of atoms involved in transforming one ligand into another. Our findings provide a key set of recommendations that should be adopted for the reliable application of RBFE methods.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Benchmarking DFT Accuracy in Predicting O 1s Binding Energies on Metals

X-ray photoelectron spectroscopy (XPS) is a powerful tool for probing the electronic structure and composition of materials, particularly metals and metal oxides of relevance to solar cells and catalysis. Density functional theory (DFT) is often used to support XPS peak assignments, but its reliability for predicting oxygen species is not well established. Here, we compile a large data set of experimental oxygen binding energies and evaluate corresponding DFT predictions. We find that as the binding energies of metal-bound atomic oxygen species increase, especially above ≈530 eV, there is a general decrease in the accuracy of DFTpredicted values. Thus, high-binding-energy atomic oxygen species, such as those proposed as active for selective Ag-catalyzed epoxidation, are less well represented. The chemical nature of the oxygen species also influences accuracy, with molecularly bound species more reliably captured across the entire range of energies. These findings illustrate the limitations of DFT for interpreting XPS spectra and provide a benchmark for improving computational methods.

Adsorption↗

Deep Learning Enabled Strain Mapping of Single-Atom Defects in Two-Dimensional Transition Metal Dichalcogenides with Sub-Picometer Precision

Two-dimensional (2D) materials offer an ideal platform to study the strain fields induced by individual atomic defects, yet challenges associated with radiation damage have so far limited electron microscopy methods to probe these atomic-scale strain fields. In this work, we demonstrate an approach to probe single-atom defects with sub-picometer precision in a monolayer 2D transition metal dichalcogenide, WSe 2–2x Te 2x . We utilize deep learning to mine large data sets of aberration-corrected scanning transmission electron microscopy images to locate and classify point defects. By combining hundreds of images of nominally identical defects, we generate high signal-to-noise class averages which allow us to measure 2D atomic spacings with up to 0.2 pm precision. Our methods reveal that Se vacancies introduce complex, oscillating strain fields in the WSe 2–2x Te 2x lattice that correspond to alternating rings of lattice expansion and contraction. These results indicate the potential impact of computer vision for the development of high-precision electron microscopy methods for beam-sensitive materials.

2D materials↗

Navigating the Expansive Landscapes of Soft Materials: A User Guide for High-Throughput Workflows

Synthetic polymers are highly customizable with tailored structures and functionality, yet this versatility generates challenges in the design of advanced materials due to the size and complexity of the design space. Thus, exploration and optimization of polymer properties using combinatorial libraries has become increasingly common, which requires careful selection of synthetic strategies, characterization techniques, and rapid processing workflows to obtain fundamental principles from these large data sets. Herein, we provide guidelines for strategic design of macromolecule libraries and workflows to efficiently navigate these high-dimensional design spaces. We describe synthetic methods for multiple library sizes and structures as well as characterization methods to rapidly generate data sets, including tools that can be adapted from biological workflows. We further highlight relevant insights from statistics and machine learning to aid in data featurization, representation, and analysis. This Perspective acts as a “user guide” for researchers interested in leveraging high-throughput screening toward the design of multifunctional polymers and predictive modeling of structure–property relationships in soft materials.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

SeismoGen: Seismic Waveform Synthesis Using GAN With Application to Seismic Data Augmentation

Abstract Detecting earthquake arrivals within seismic time series can be a challenging task. Visual, human detection has long been considered the gold standard but requires intensive manual labor that scales poorly to large data sets. In recent years, automatic detection methods based on machine learning have been developed to improve the accuracy and efficiency. However, the accuracy of those methods relies on access to a sufficient amount of high‐quality labeled training data, often tens of thousands of records or more. We aim to resolve this dilemma by answering two questions: (1) provided with a limited amount of reliable labeled data, can we use them to generate additional, realistic synthetic waveform data? and (2) can we use those synthetic data to further enrich the training set through data augmentation, thereby enhancing detection algorithms? To address these questions, we use a generative adversarial network (GAN), a type of machine learning model which has shown supreme capability in generating high‐quality synthetic samples in multiple domains. Once trained, our GAN model is capable of producing realistic seismic waveforms of multiple labels (noise and event classes). Applied to real Earth seismic data sets in Oklahoma, we show that data augmentation from our GAN‐generated synthetic waveforms can be used to improve earthquake detection algorithms in instances when only small amounts of labeled training data are available.

Wang, Tiantong↗

Scenario Discovery Analysis of Drivers of Solar and Wind Energy Transitions Through 2050

Deep human-Earth system uncertainties and strong multi-sector dynamics make it difficult to anticipate which conditions are most likely to lead to higher or lower adoption of renewable energy, and models project a broad range of future solar and wind energy shares across future scenarios. To elucidate these dynamics, we explore a large data set of scenarios simulated from the Global Change Analysis Model (GCAM) and use scenario discovery to identify the most significant factors affecting solar and wind adoption by mid-century. We generated a data set of over 4,000 scenarios from GCAM by varying 12 different socioeconomic factors at high and low levels, including assumptions about future energy demand, resource costs, and fossil fuel emissions paths, as well as specific technology assumptions including wind and solar backup requirements and storage costs. Using scenario discovery, we assess the most important factors globally and regionally in creating high fractions of solar and wind energy and explore interconnected effects on other systems including water and non-CO 2 emissions. Globally and regionally, we found that solar and wind-related technology costs were the primary drivers of high wind and solar energy adoption, though a few regions depend heavily on other parameters like carbon capture and storage costs, population and gross domestic product trajectories, and fossil fuel costs. We also identify four key paths to high solar and wind energy by mid-century and discuss their tradeoffs in terms of other outcomes.

14 SOLAR ENERGY↗

The Collaborative Seismic Earth Model: Generation 2

Geological interpretations, earthquake source inversions and ground motion modeling, among other applications, require models that jointly resolve crustal and mantle structure. With the second generation of the Collaborative Seismic Earth Model (CSEM2), we present a global multi-resolution tomographic Earth model that serves this purpose. The model evolves through successive regional- and global-scale refinements. While the first generation aggregated regional models, with this study, we ensure consistency between all individual submodels, resulting in a model that accurately explains wave propagation across scales. Recent regional tomographic models were incorporated, comprising continental-scale inversions for Asia and Africa, as well as regional inversions for the Western US, Central Andes, Iran, and Southeast Asia. Across all regional refinements, over 793,000 source-receiver pairs contributed. Moreover, the long-wavelength Earth model (LOWE) introduces large-scale structures outside of pre-existing local refinements. A full-waveform inversion for global anisotropic P-and S-wave speed structure over a total of 194 iterations with a minimum period of 50 s on a large data set of 1 hr of waveform data from 2,423 earthquakes and over 6 million source-receiver pairs ensures that regional updates in the crust and uppermost mantle translate into updates of deeper, global-scale structure. To test the performance of CSEM2, we evaluate waveform fits between observed and synthetic seismograms at 50 s for an independent data set on the global scale, and on the regional scale for lower periods. We accurately simulate waveforms within and across regional refinements, maintaining the original resolution of the submodels embedded in the global framework.

58 GEOSCIENCES↗

Leveraging generative adversarial networks to create realistic scanning transmission electron microscopy images

Abstract The rise of automation and machine learning (ML) in electron microscopy has the potential to revolutionize materials research through autonomous data collection and processing. A significant challenge lies in developing ML models that rapidly generalize to large data sets under varying experimental conditions. We address this by employing a cycle generative adversarial network (CycleGAN) with a reciprocal space discriminator, which augments simulated data with realistic spatial frequency information. This allows the CycleGAN to generate images nearly indistinguishable from real data and provide labels for ML applications. We showcase our approach by training a fully convolutional network (FCN) to identify single atom defects in a 4.5 million atom data set, collected using automated acquisition in an aberration-corrected scanning transmission electron microscope (STEM). Our method produces adaptable FCNs that can adjust to dynamically changing experimental variables with minimal intervention, marking a crucial step towards fully autonomous harnessing of microscopy big data.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

Machine learning inversion from small-angle scattering for charged polymers

We develop Monte Carlo simulations for uniformly charged polymers and a machine learning algorithm to interpret the intra-polymer structure factor of the charged polymer system, which can be obtained from small-angle scattering experiments. The polymer is modeled as a chain of fixed-length bonds, where the connected bonds are subject to bending energy, and there is also a screened Coulomb potential for charge interaction between all joints. The bending energy is determined by the intrinsic bending stiffness, and the charge interaction depends on the interaction strength and screening length. All three contribute to the stiffness of the polymer chain and lead to longer and larger polymer conformations. The screening length also introduces a second length scale for the polymer besides the bending persistence length. To obtain the inverse mapping from the structure factor to these polymer conformation and energy-related parameters, we generate a large data set of structure factors by running simulations for a wide range of polymer energy parameters. We use principal component analysis to investigate the intra-polymer structure factors and determine the feasibility of the inversion using the nearest neighbor distance. We employ Gaussian process regression to achieve the inverse mapping and extract the characteristic parameters of polymers from the structure factor with low relative error.

36 MATERIALS SCIENCE↗

Estimates of the wavenumber wavelet power spectrum of magnetic fluctuations during magnetic reconnection

Fluctuation analyses of experimental observations generally lack high temporal resolution and are in frequency-space f, contrary to theoretical efforts in wavenumber-space k. This is due to the inherent limits of the Fourier transform, though it is prominent due to the ease of diagnostic implementation. Advances in wavelet-based analysis have provided relief due to its temporal resolution, but in its common use, is still hard to compare to theoretical models. By using the two-point correlation technique in conjunction with large data sets, a wavelet power spectrum in wavenumber-space can be created. Dubbed the wavenumber wavelet power spectrum, this spectrum relates wavenumber to power in time. Further, this analysis technique more closely connects characterizations of experimentally observed fluctuations with other system parameters and theoretical predictions. In this article, we develop the wavenumber wavelet power spectrum using magnetic fluctuations caused by tearing instability driven magnetic reconnection in reproducible, high temperature laboratory plasmas. These dynamic magnetic fluctuations generated in reversed field pinch plasmas are broadband, ranging from the low frequency, 10's of kHz, up to the ion gyroradii frequencies, 100's of kHz. The dominant fluctuations have poloidal and toroidal mode numbers (m,n)=(1,6−10) and can grow to 2%–3% of the mean magnetic field. During these reconnection events, ions, and electrons are energized, magnetic fluctuation amplitudes increase, plasma flow is halted, and the toroidal magnetic flux increases, all on a semi-periodic basis. The newly developed spectrum provides better temporal resolution of spectrum characteristics to correlate with these particle energization phenomena.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Identification of low-momentum muons in the CMS detector using multivariate techniques in proton-proton collisions at $\sqrt{s}$ = 13.6 TeV

“Soft” muons with a transverse momentum below 10 GeV are featured in many processes studied by the CMS experiment, such as decays of heavy-flavor hadrons or rare tau lepton decays. Maximizing the selection efficiency for these muons, while simultaneously suppressing backgrounds from long-lived light-flavor hadron decays, is therefore important for the success of the CMS physics program. Multivariate techniques have been shown to deliver better muon identification performance than traditional selection techniques. To take full advantage of the large data set currently being collected during Run 3 of the CERN LHC, a new multivariate classifier based on a gradient-boosted decision tree has been developed. It offers a significantly improved separation of signal and background muons compared to a similar classifier used for the analysis of the Run 2 data. The performance of the new classifier is evaluated on a data set collected with the CMS detector in 2022 and 2023, corresponding to an integrated luminosity of 62 fb -1 .

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Genetic and Functional Diversity Help Explain Pathogenic, Weakly Pathogenic, and Commensal Lifestyles in the Genus Xanthomonas

The genus Xanthomonas has been primarily studied for pathogenic interactions with plants. However, besides host and tissue-specific pathogenic strains, this genus also comprises nonpathogenic strains isolated from a broad range of hosts, sometimes in association with pathogenic strains, and other environments, including rainwater. Based on their incapacity or limited capacity to cause symptoms on the host of isolation, nonpathogenic xanthomonads can be further characterized as commensal and weakly pathogenic. This study aimed to understand the diversity and evolution of nonpathogenic xanthomonads compared to their pathogenic counterparts based on their cooccurrence and phylogenetic relationship and to identify genomic traits that form the basis of a life history framework that groups xanthomonads by ecological strategies. We sequenced genomes of 83 strains spanning the genus phylogeny and identified eight novel species, indicating unexplored diversity. While some nonpathogenic species have experienced a recent loss of a type III secretion system, specifically the hrp2 cluster, we observed an apparent lack of association of the hrp2 cluster with lifestyles of diverse species. We performed association analysis on a large data set of 337 Xanthomonas strains to explain how xanthomonads may have established association with the plants across the continuum of lifestyles from commensals to weak pathogens to pathogens. Presence of distinct transcriptional regulators, distinct nutrient utilization and assimilation genes, transcriptional regulators, and chemotaxis genes may explain lifestyle-specific adaptations of xanthomonads.

59 BASIC BIOLOGICAL SCIENCES↗

Moment tensor event identification for collapses

SUMMARY We introduce a seismic identification method for collapse events using moment tensors (MTs). We start by computing full (six-element) MT solutions for 43 identified collapse events from around the world, and statistically characterizing the population on the MT hypersphere. We then test a large data set of over 1000 full MTs for the western U.S. against the distribution of collapses using a MT-based identification method similarly as used for testing explosions. Known collapses and explosions are readily identified, along with other anomalous events in the Geysers and central California coast. Misidentification rates are determined for various screening angles with optimal misidentification rates between earthquakes and collapses on the order of 3 per cent. The method is demonstrated to be very effective at identifying non-earthquake sources with a 97–98 per cent accuracy. It is likely to be transportable to other regions, and can be used for event identification anywhere full MT solutions are routinely calculated.

58 GEOSCIENCES↗

A unique, ring-like radio source with quadrilateral structure detected with machine learning

ABSTRACT We report the discovery of a unique object in the MeerKAT Galaxy Cluster Legacy Survey (MGCLS) using the machine learning anomaly detection framework astronomaly. This strange, ring-like source is 30′ from the MGCLS field centred on Abell 209, and is not readily explained by simple physical models. With an assumed host galaxy at redshift 0.55, the luminosity (1025 W Hz−1) is comparable to powerful radio galaxies. The source consists of a ring of emission 175 kpc across, quadrilateral enhanced brightness regions bearing resemblance to radio jets, two ‘ears’ separated by 368 kpc, and a diffuse envelope. All of the structures appear spectrally steep, ranging from −1.0 to −1.5. The ring has high polarization (25 per cent) except on the bright patches (<10 per cent). We compare this source to the Odd Radio Circles recently discovered in ASKAP data and discuss several possible physical models, including a termination shock from starburst activity, an end-on radio galaxy, and a supermassive black hole merger event. No simple model can easily explain the observed structure of the source. This work, as well as other recent discoveries, demonstrates the power of unsupervised machine learning in mining large data sets for scientifically interesting sources.

Astronomy & Astrophysics↗