Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “large data sets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

A riming‐dependent parameterization of scattering by snowflakes using the self‐similar Rayleigh–Gans approximation

Abstract Riming is a key process of precipitation formation in ice‐containing clouds, but quantifying riming from observations is challenging, limiting our ability to evaluate the riming process in numerical weather models. One challenge for radar observations is that riming changes both the physical properties (mass, area cross‐section) and scattering properties of ice particles. These changes need to be implemented consistently as a function of riming in radar forward operators, which are required for retrievals and model evaluation in observation space. In this study, mass–size, cross‐section area–size, and backscattering cross‐section relations are developed as a function of the normalized rime mass for aggregates composed of various monomer types (columns, dendrites, needles, plates, and rosettes). The proposed framework allows us to simulate scattering properties of aggregated ice particles consistently as a function of riming in retrievals and radar forward operators. The parameterizations are developed from a large data set of simulated rimed aggregates of different sizes and monomer crystal types. The backscattering cross‐section parameterization (the “riming‐dependent parameterization”) is evaluated for radar frequencies of 35.6 and 94.0 GHz and is based on the Self‐Similar Rayleigh–Gans approximation (SSRGA), which is increasingly used to calculate microwave scattering of ice crystals and snowflakes. Compared with parameterizations from the literature that do not consider riming, the riming‐dependent parameterization leads to significantly smaller biases in terms of backscattering cross‐section. When using the particle masses and scattering properties of the individual particles simulated by the aggregation and riming model as a reference, the bias of our parameterization is below 1 dB when integrating over an exponential particle size distribution with sizes from 0.1–10 mm.

54 ENVIRONMENTAL SCIENCES↗

libwfa: Wavefunction analysis tools for excited and open‐shell electronic states

Abstract An open‐source software library for wavefunction analysis, libwfa, provides a comprehensive and flexible toolbox for post‐processing excited‐state calculations, featuring a hierarchy of interconnected visual and quantitative analysis methods. These tools afford compact graphical representations of various excited‐state processes, provide detailed insight into electronic structure, and are suitable for automated processing of large data sets. The analysis is based on reduced quantities, such as state and transition density matrices (DMs), and allows one to distill simple molecular orbital pictures of physical phenomena from intricate correlated wavefunctions. The implemented descriptors provide a rigorous link between many‐body wavefunctions and intuitive physical and chemical models, for example, exciton binding, double excitations, orbital relaxation, and polyradical character. A broad range of quantum‐chemical methods is interfaced with libwfa via a uniform interface layer in the form of DMs. This contribution reviews the structure of libwfa and highlights its capabilities by several representative use cases. This article is categorized under: Software > Quantum Chemistry Theoretical and Physical Chemistry > Spectroscopy

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

NanoSIP: NanoSIMS Applications for Microbial Biology

High-resolution imaging with secondary ion mass spectrometry (nanoSIMS) has become a standard method in systems biology and environmental biogeochemistry and is broadly used to decipher ecophysiological traits of environmental microorganisms, metabolic processes in plant and animal tissues, and cross-kingdom symbioses. When combined with stable isotope-labeling—an approach we refer to as nanoSIP—nanoSIMS imaging offers a distinctive means to quantify net assimilation rates and stoichiometry of individual cell-sized particles in both low- and high-complexity environments. While the majority of nanoSIP studies in environmental and microbial biology have focused on nitrogen and carbon metabolism (using 15 N and 13 C tracers), multiple advances have pushed the capabilities of this approach in the past decade. The development of a high-brightness oxygen ion source has enabled high-resolution metal analyses that are easier to perform, allowing quantification of metal distribution in cells and environmental particles. New preparation methods, tools for automated data extraction from large data sets, and analytical approaches that push the limits of sensitivity and spatial resolution have allowed for more robust characterization of populations ranging from marine archaea to fungi and viruses. Further, NanoSIMS studies continue to be enhanced by correlation with orthogonal imaging and ‘omics approaches; when linked to molecular visualization methods, such as in situ hybridization and antibody labeling, these techniques enable in situ function to be linked to microbial identity and gene expression. Here we present an updated description of the primary materials, methods, and calculations used for nanoSIP, with an emphasis on recent advances in nanoSIMS applications, key methodological steps, and potential pitfalls.

07 ISOTOPE AND RADIATION SOURCES↗

Interactive Exploration of High-Dimensional Phase Diagrams

High-dimensional thermodynamic phase stability databases are becoming increasingly common due to the convergence of three recent trends: (i) the widespread interest in so-called “high-entropy” alloys, (ii) the availability of high-throughput computational assessments of phase stability in broad composition spaces and (iii) the ongoing development of ever-increasingly broad, multicomponent, multiphase CALPHAD databases. Although automated computational tools can readily process such high-dimensional data, scientists are often unable to visualize the relevant phase relations, an ability that is crucial to gaining an intuitive understanding of the stability constraints governing materials design. The present work addresses this need by providing algorithms that enable the interactive exploration of phase equilibria in high-dimensional spaces. These algorithms concentrate the complex nonlinear nonsmooth optimization needed into a preprocessing step that generates a large number of high-dimensional yet elementary graphical primitives. Furthermore, these primitives can then be cross-sectioned to yield 3-dimensional views in a computationally efficient manner that enables an interactive exploration of high-dimensional spaces. All of these operations are highly parallelizable, thus facilitating scaling of this method to large data sets.

36 MATERIALS SCIENCE↗

Piecewise linear approximation with minimum number of linear segments and minimum error: A fast approach to tighten and warm start the hierarchical mixed integer formulation

In several areas of economics and engineering, it is often necessary to fit discrete data points or approximate nonlinear functions with continuous functions. Piecewise linear (PWL) functions are a convenient way to achieve this. PWL functions can be modeled in mathematical problems using only linear and integer variables. Moreover, there is a computational benefit in using PWL functions that have the least possible number of segments. This work proposes a novel hierarchical mixed integer linear programming (MILP) formulation that identifies a continuous PWL approximation with minimum number of linear segments for a given target maximum error. The proposed MILP formulation also identifies the solution with the least maximum error among the solutions with minimum number of segments. Then, this work proposes a fast iterative algorithm that identifies non necessarily continuous PWL approximations by solving O(S log N) linear programming (LP) problems, where N is the number of data points and S is the minimum number of segments in the non necessarily continuous case. This work demonstrates that tight bounds for the MILP problem can be derived from these approximations. Next, a fast algorithm is introduced to transform a non necessarily continuous PWL approximation into a continuous one. Finally, the tight bounds and the continuous PWL approximations are used to tighten and warm start the MILP problem. The tightened formulation is shown in experimental results to be more efficient, especially for large data sets, with a solution time that is up to two orders of magnitude less than the existing literature.

97 MATHEMATICS AND COMPUTING↗

In-depth analysis on parallel processing patterns for high-performance Dataframes

The Data Science domain has expanded monumentally in both research and industry communities during the past decade, predominantly owing to the Big Data revolution. Artificial Intelligence (AI) and Machine Learning (ML) are bringing more complexities to data engineering applications, which are now integrated into data processing pipelines to process terabytes of data. Typically, a significant amount of time is spent on data preprocessing in these pipelines, and hence improving its efficiency directly impacts the overall pipeline performance. The community has recently embraced the concept of Dataframes as the de-facto data structure for data representation and manipulation. However, the most widely used serial Dataframes today (R, pandas) experience performance limitations while working on even moderately large data sets. We believe that there is plenty of room for improvement by taking a look at this problem from a high-performance computing point of view. In a prior publication, we presented a set of parallel processing patterns for distributed dataframe operators and the reference runtime implementation, Cylon. In this paper, we are expanding on the initial concept by introducing a cost model for evaluating the said patterns. Furthermore, we evaluate the performance of Cylon on the ORNL Summit supercomputer.

97 MATHEMATICS AND COMPUTING↗

Solving differential equations using deep neural networks

Recent work on solving partial differential equations (PDEs) with deep neural networks (DNNs) is presented. The paper reviews and extends some of these methods while carefully analyzing a fundamental feature in numerical PDEs and nonlinear analysis: irregular solutions. First, the Sod shock tube solution to the compressible Euler equations is discussed and analyzed. This analysis includes a comparison of a DNN-based approach with conventional finite element and finite volume methods, and demonstrates that the DNN is competitive in terms of degrees of freedom required for a given accuracy. Further, the DNN-based approach is extended to consider performance improvements and simultaneous parameter space exploration. Next, a shock solution to compressible magnetohydrodynamics (MHD) is solved for, and used in a scenario where experimental data is utilized to enhance a PDE system that is a priori insufficient to validate against the observed/experimental data. This is accomplished by enriching the model PDE system with source terms that are then inferred via supervised training with synthetic experimental data. The resulting DNN framework for PDEs enables straightforward system prototyping and natural integration of large data sets (be they synthetic or experimental), all while simultaneously enabling single-pass exploration of an entire parameter space.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Investigating the Role of Copper in Arsenic Doped Cd(Se,Te) Photovoltaics

The open circuit voltage (VOC) deficit in Cd(Se,Te)-based photovoltaics remains a critical obstacle for pushing the technology closer to theoretical performance limits. Arsenic doping has become a dominant and promising route to achieve the higher p-type carrier concentrations necessary for higher VOC, but challenges associated with this alternate defect chemistry and higher doping density have hindered progress. Here we show that while arsenic doping enables high carrier concentrations (>1016 cm-3), co-doping with copper can provide a boost to VOC without a significant change to carrier concentration. A large data set is initially used to explore current-voltage and capacitance-voltage trends associated with arsenic doped devices with and without copper. A smaller subset is then used to probe these trends using a wide variety of characterization techniques. Copper is found to facilitate reduced interface recombination and potentially improved bulk absorber characteristics, though the mechanisms for these improvements are not yet clear. Despite the improved performance of co-doped devices, VOC is still far below its potential especially for highly doped devices. Low emitter doping in conjunction with high absorber doping seems to be a plausible cause for this significant deficit, though other device properties may exacerbate this problem.

CdSeTe↗

Deep Learning-Enabled MS/MS Spectrum Prediction Facilitates Automated Identification Of Novel Psychoactive Substances

The market for illicit drugs has been reshaped by the emergence of more than 1100 new psychoactive substances (NPS) over the past decade, posing a major challenge to the forensic and toxicological laboratories tasked with detecting and identifying them. Tandem mass spectrometry (MS/MS) is the primary method used to screen for NPS within seized materials or biological samples. The most contemporary workflows necessitate labor-intensive and expensive MS/MS reference standards, which may not be available for recently emerged NPS on the illicit market. Here, we present NPS-MS, a deep learning method capable of accurately predicting the MS/MS spectra of known and hypothesized NPS from their chemical structures alone. NPS-MS is trained by transfer learning from a generic MS/MS prediction model on a large data set of MS/MS spectra. We show that this approach enables a more accurate identification of NPS from experimentally acquired MS/MS spectra than any existing method. We demonstrate the application of NPS-MS to identify a novel derivative of phencyclidine (PCP) within an unknown powder seized in Denmark without the use of any reference standards. We anticipate that NPS-MS will allow forensic laboratories to identify more rapidly both known and newly emerging NPS. NPS-MS is available as a web server at https://nps-ms.ca/, which provides MS/MS spectra prediction capabilities for given NPS compounds. Additionally, it offers MS/MS spectra identification against a vast database comprising approximately 8.7 million predicted NPS compounds from DarkNPS and 24.5 million predicted ESI-QToF-MS/MS spectra for these compounds.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Predicting Small Molecule Transfer Free Energies by Combining Molecular Dynamics Simulations and Deep Learning

Accurately predicting small molecule partitioning and hydrophobicity is critical in the drug discovery process. There are many heterogeneous chemical environments within a cell and entire human body. For example, drugs must be able to cross the hydrophobic cellular membrane to reach their intracellular targets, and hydrophobicity is an important driving force for drug–protein binding. Atomistic molecular dynamics (MD) simulations are routinely used to calculate free energies of small molecules binding to proteins, crossing lipid membranes, and solvation but are computationally expensive. Machine learning (ML) and empirical methods are also used throughout drug discovery but rely on experimental data, limiting the domain of applicability. We present atomistic MD simulations calculating 15,000 small molecule free energies of transfer from water to cyclohexane. This large data set is used to train ML models that predict the free energies of transfer. We show that a spatial graph neural network model achieves the highest accuracy, followed closely by a 3D-convolutional neural network, and shallow learning based on the chemical fingerprint is significantly less accurate. A mean absolute error of ~4 kJ/mol compared to the MD calculations was achieved for our best ML model. We also show that including data from the MD simulation improves the predictions, tests the transferability of each model to a diverse set of molecules, and show multitask learning improves the predictions. This work provides insight into the hydrophobicity of small molecules and ML cheminformatics modeling, and our data set will be useful for designing and testing future ML cheminformatics methods.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Large Scale Study of Ligand–Protein Relative Binding Free Energy Calculations: Actionable Predictions from Statistically Robust Protocols

The accurate and reliable prediction of protein–ligand binding affinities can play a central role in the drug discovery process as well as in personalized medicine. Of considerable importance during lead optimization are the alchemical free energy methods that furnish an estimation of relative binding free energies (RBFE) of similar molecules. Recent advances in these methods have increased their speed, accuracy, and precision. This is evident from the increasing number of retrospective as well as prospective studies employing them. However, such methods still have limited applicability in real-world scenarios due to a number of important yet unresolved issues. Here, we report the findings from a large data set comprising over 500 ligand transformations spanning over 300 ligands binding to a diverse set of 14 different protein targets which furnish statistically robust results on the accuracy, precision, and reproducibility of RBFE calculations. We use ensemble-based methods which are the only way to provide reliable uncertainty quantification given that the underlying molecular dynamics is chaotic. These are implemented using TIES (Thermodynamic Integration with Enhanced Sampling). Results achieve chemical accuracy in all cases. Ensemble simulations also furnish information on the statistical distributions of the free energy calculations which exhibit non-normal behavior. We find that the “enhanced sampling” method known as replica exchange with solute tempering degrades RBFE predictions. We also report definitively on numerous associated alchemical factors including the choice of ligand charge method, flexibility in ligand structure, and the size of the alchemical region including the number of atoms involved in transforming one ligand into another. Our findings provide a key set of recommendations that should be adopted for the reliable application of RBFE methods.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Benchmarking DFT Accuracy in Predicting O 1s Binding Energies on Metals

X-ray photoelectron spectroscopy (XPS) is a powerful tool for probing the electronic structure and composition of materials, particularly metals and metal oxides of relevance to solar cells and catalysis. Density functional theory (DFT) is often used to support XPS peak assignments, but its reliability for predicting oxygen species is not well established. Here, we compile a large data set of experimental oxygen binding energies and evaluate corresponding DFT predictions. We find that as the binding energies of metal-bound atomic oxygen species increase, especially above ≈530 eV, there is a general decrease in the accuracy of DFTpredicted values. Thus, high-binding-energy atomic oxygen species, such as those proposed as active for selective Ag-catalyzed epoxidation, are less well represented. The chemical nature of the oxygen species also influences accuracy, with molecularly bound species more reliably captured across the entire range of energies. These findings illustrate the limitations of DFT for interpreting XPS spectra and provide a benchmark for improving computational methods.

Adsorption↗

Deep Learning Enabled Strain Mapping of Single-Atom Defects in Two-Dimensional Transition Metal Dichalcogenides with Sub-Picometer Precision

Two-dimensional (2D) materials offer an ideal platform to study the strain fields induced by individual atomic defects, yet challenges associated with radiation damage have so far limited electron microscopy methods to probe these atomic-scale strain fields. In this work, we demonstrate an approach to probe single-atom defects with sub-picometer precision in a monolayer 2D transition metal dichalcogenide, WSe 2–2x Te 2x . We utilize deep learning to mine large data sets of aberration-corrected scanning transmission electron microscopy images to locate and classify point defects. By combining hundreds of images of nominally identical defects, we generate high signal-to-noise class averages which allow us to measure 2D atomic spacings with up to 0.2 pm precision. Our methods reveal that Se vacancies introduce complex, oscillating strain fields in the WSe 2–2x Te 2x lattice that correspond to alternating rings of lattice expansion and contraction. These results indicate the potential impact of computer vision for the development of high-precision electron microscopy methods for beam-sensitive materials.

2D materials↗

Navigating the Expansive Landscapes of Soft Materials: A User Guide for High-Throughput Workflows

Synthetic polymers are highly customizable with tailored structures and functionality, yet this versatility generates challenges in the design of advanced materials due to the size and complexity of the design space. Thus, exploration and optimization of polymer properties using combinatorial libraries has become increasingly common, which requires careful selection of synthetic strategies, characterization techniques, and rapid processing workflows to obtain fundamental principles from these large data sets. Herein, we provide guidelines for strategic design of macromolecule libraries and workflows to efficiently navigate these high-dimensional design spaces. We describe synthetic methods for multiple library sizes and structures as well as characterization methods to rapidly generate data sets, including tools that can be adapted from biological workflows. We further highlight relevant insights from statistics and machine learning to aid in data featurization, representation, and analysis. This Perspective acts as a “user guide” for researchers interested in leveraging high-throughput screening toward the design of multifunctional polymers and predictive modeling of structure–property relationships in soft materials.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

SeismoGen: Seismic Waveform Synthesis Using GAN With Application to Seismic Data Augmentation

Abstract Detecting earthquake arrivals within seismic time series can be a challenging task. Visual, human detection has long been considered the gold standard but requires intensive manual labor that scales poorly to large data sets. In recent years, automatic detection methods based on machine learning have been developed to improve the accuracy and efficiency. However, the accuracy of those methods relies on access to a sufficient amount of high‐quality labeled training data, often tens of thousands of records or more. We aim to resolve this dilemma by answering two questions: (1) provided with a limited amount of reliable labeled data, can we use them to generate additional, realistic synthetic waveform data? and (2) can we use those synthetic data to further enrich the training set through data augmentation, thereby enhancing detection algorithms? To address these questions, we use a generative adversarial network (GAN), a type of machine learning model which has shown supreme capability in generating high‐quality synthetic samples in multiple domains. Once trained, our GAN model is capable of producing realistic seismic waveforms of multiple labels (noise and event classes). Applied to real Earth seismic data sets in Oklahoma, we show that data augmentation from our GAN‐generated synthetic waveforms can be used to improve earthquake detection algorithms in instances when only small amounts of labeled training data are available.

Wang, Tiantong↗

Scenario Discovery Analysis of Drivers of Solar and Wind Energy Transitions Through 2050

Deep human-Earth system uncertainties and strong multi-sector dynamics make it difficult to anticipate which conditions are most likely to lead to higher or lower adoption of renewable energy, and models project a broad range of future solar and wind energy shares across future scenarios. To elucidate these dynamics, we explore a large data set of scenarios simulated from the Global Change Analysis Model (GCAM) and use scenario discovery to identify the most significant factors affecting solar and wind adoption by mid-century. We generated a data set of over 4,000 scenarios from GCAM by varying 12 different socioeconomic factors at high and low levels, including assumptions about future energy demand, resource costs, and fossil fuel emissions paths, as well as specific technology assumptions including wind and solar backup requirements and storage costs. Using scenario discovery, we assess the most important factors globally and regionally in creating high fractions of solar and wind energy and explore interconnected effects on other systems including water and non-CO 2 emissions. Globally and regionally, we found that solar and wind-related technology costs were the primary drivers of high wind and solar energy adoption, though a few regions depend heavily on other parameters like carbon capture and storage costs, population and gross domestic product trajectories, and fossil fuel costs. We also identify four key paths to high solar and wind energy by mid-century and discuss their tradeoffs in terms of other outcomes.

14 SOLAR ENERGY↗

The Collaborative Seismic Earth Model: Generation 2

Geological interpretations, earthquake source inversions and ground motion modeling, among other applications, require models that jointly resolve crustal and mantle structure. With the second generation of the Collaborative Seismic Earth Model (CSEM2), we present a global multi-resolution tomographic Earth model that serves this purpose. The model evolves through successive regional- and global-scale refinements. While the first generation aggregated regional models, with this study, we ensure consistency between all individual submodels, resulting in a model that accurately explains wave propagation across scales. Recent regional tomographic models were incorporated, comprising continental-scale inversions for Asia and Africa, as well as regional inversions for the Western US, Central Andes, Iran, and Southeast Asia. Across all regional refinements, over 793,000 source-receiver pairs contributed. Moreover, the long-wavelength Earth model (LOWE) introduces large-scale structures outside of pre-existing local refinements. A full-waveform inversion for global anisotropic P-and S-wave speed structure over a total of 194 iterations with a minimum period of 50 s on a large data set of 1 hr of waveform data from 2,423 earthquakes and over 6 million source-receiver pairs ensures that regional updates in the crust and uppermost mantle translate into updates of deeper, global-scale structure. To test the performance of CSEM2, we evaluate waveform fits between observed and synthetic seismograms at 50 s for an independent data set on the global scale, and on the regional scale for lower periods. We accurately simulate waveforms within and across regional refinements, maintaining the original resolution of the submodels embedded in the global framework.

58 GEOSCIENCES↗

Leveraging generative adversarial networks to create realistic scanning transmission electron microscopy images

Abstract The rise of automation and machine learning (ML) in electron microscopy has the potential to revolutionize materials research through autonomous data collection and processing. A significant challenge lies in developing ML models that rapidly generalize to large data sets under varying experimental conditions. We address this by employing a cycle generative adversarial network (CycleGAN) with a reciprocal space discriminator, which augments simulated data with realistic spatial frequency information. This allows the CycleGAN to generate images nearly indistinguishable from real data and provide labels for ML applications. We showcase our approach by training a fully convolutional network (FCN) to identify single atom defects in a 4.5 million atom data set, collected using automated acquisition in an aberration-corrected scanning transmission electron microscope (STEM). Our method produces adaptable FCNs that can adjust to dynamically changing experimental variables with minimal intervention, marking a crucial step towards fully autonomous harnessing of microscopy big data.

77 NANOSCIENCE AND NANOTECHNOLOGY↗