Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “large data sets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Organizing Large Data Sets for Efficient Analyses on HPC Systems

Upcoming exascale applications could introduce significant data management challenges due to their large sizes, dynamic work distribution, and involvement of accelerators such as graphical processing units, GPUs. In this work, we explore the performance of reading and writing operations involving one such scientific application on two different supercomputers. Our tests showed that the Adaptable Input and Output System, ADIOS, was able to achieve speeds over 1TB/s, a significant fraction of the peak I/O performance on Summit. We also demonstrated the querying functionality in ADIOS could effectively support common selective data analysis operations, such as conditional histograms. In tests, this query mechanism was able to reduce the execution time by a factor of five. More importantly, ADIOS data management framework allows us to achieve these performance improvements with only a minimal amount of coding effort.

Gu, Junmin↗

The Rise of Neural Networks for Materials and Chemical Dynamics

Machine learning (ML) is quickly becoming a premier tool for modeling chemical processes and materials. ML-based force fields, trained on large data sets of high-quality electron structure calculations, are particularly attractive due their unique combination of computational efficiency and physical accuracy. This Perspective summarizes some recent advances in the development of neural network-based interatomic potentials. Designing high-quality training data sets is crucial to overall model accuracy. One strategy is active learning, in which new data are automatically collected for atomic configurations that produce large ML uncertainties. Another strategy is to use the highest levels of quantum theory possible. Transfer learning allows training to a data set of mixed fidelity. A model initially trained to a large data set of density functional theory calculations can be significantly improved by retraining to a relatively small data set of expensive coupled cluster theory calculations. These advances are exemplified by applications to molecules and materials.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Crack detection in fuel cell electrodes using a spatial filtering technique for overcoming noisy backgrounds

Image processing is a powerful tool that allows for rapid and automated data parsing in settings that occupy large variable spaces and require large data sets. Feature detection on difficultly discerned backgrounds is a subset of image processing that facilitates the extraction of quantitative metrics from otherwise subjective data. Crack detection and quantification is an important capability in polymer electrolyte membrane fuel cell quality control, failure analysis, and optimization. This work presents a technique to perform crack detection and quantification which overcomes challenges faced by commonly used image segmentation techniques. We demonstrate the use of a geometrically filtered noise‐level detection technique to select a binary threshold value from which we then quantify how cracked a sample is. Furthermore, we demonstrate the accuracy of our technique using programmatically generated test images of known crack amounts and their performance on real‐world fuel cell catalyst layer samples.

30 DIRECT ENERGY CONVERSION↗

Quantifying Atomically Dispersed Catalysts Using Deep Learning Assisted Microscopy

The catalytic performance of atomically dispersed catalysts (ADCs) is greatly influenced by their atomic configurations, such as atom–atom distances, clustering of atoms into dimers and trimers, and their distributions. Scanning transmission electron microscopy (STEM) is a powerful technique for imaging ADCs at the atomic scale; however, most STEM analyses of ADCs thus far have relied on human labeling, making it difficult to analyze large data sets. Here, we introduce a convolutional neural network (CNN)-based algorithm capable of quantifying the spatial arrangement of different adatom configurations. The algorithm was tested on different ADCs with varying support crystallinity and homogeneity. Results show that our algorithm can accurately identify atom positions and effectively analyze large data sets. Here, this work provides a robust method to overcome a major bottleneck in STEM analysis for ADC catalyst research. We highlight the potential of this method to serve as an on-the-fly analysis tool for catalysts in future in situ microscopy experiments.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Radiative properties of quantum emitters in boron nitride from excited state calculations and Bayesian analysis

Abstract Point defects in hexagonal boron nitride (hBN) have attracted growing attention as bright single-photon emitters. However, understanding of their atomic structure and radiative properties remains incomplete. Here we study the excited states and radiative lifetimes of over 20 native defects and carbon or oxygen impurities in hBN using ab initio density functional theory and GW plus Bethe-Salpeter equation calculations, generating a large data set of their emission energy, polarization and lifetime. We find a wide variability across quantum emitters, with exciton energies ranging from 0.3 to 4 eV and radiative lifetimes from ns to ms for different defect structures. Through a Bayesian statistical analysis, we identify various high-likelihood charge-neutral defect emitters, among which the native V N N B defect is predicted to possess emission energy and radiative lifetime in agreement with experiments. Our work advances the microscopic understanding of hBN single-photon emitters and introduces a computational framework to characterize and identify quantum emitters in 2D materials.

Chemistry↗

The System for Classification of Low-Pressure Systems (SyCLoPS): An All-In-One Objective Framework for Large-Scale Data Sets

We propose the first unified objective framework (SyCLoPS) for detecting and classifying all types of low-pressure systems (LPSs) in a given data set. We use the state-of-the-art automated feature tracking software TempestExtremes (TE) to detect and track LPS features globally in ERA5 and compute 16 parameters from commonly found atmospheric variables for classification. A Python classifier is implemented to classify all LPSs at once. The framework assigns 16 different labels (classes) to each LPS data point and designates four different types of high-impact LPS tracks, including tracks of tropical cyclone (TC), monsoonal system, subtropical storm and polar low. The classification process involves disentangling high-altitude and drier LPSs, differentiating tropical and non-tropical LPSs using novel criteria, and optimizing for the detection of the four types of high-impact LPS. A comparison of our labels with those in the International Best Track Archive for Climate Stewardship (IBTrACS) revealed an overall accuracy of 95% in distinguishing between tropical systems, extratropical cyclones, and disturbances. SyCLoPS produces a better TC detection skill compared to the previous algorithms, highlighted by an approximately 6% reduction in the false alarm rate compared to the previous TE algorithm. The vertical cross section composite of the four types of high-impact LPS we detect each shows distinct structural characteristics. Finally, we demonstrate that SyCLoPS is valuable for investigating various aspects of LPSs in climate data, such as the evolution of a single LPS track, patterns of LPS frequencies, and precipitation or wind influence associated with a particular LPS class.

54 ENVIRONMENTAL SCIENCES↗

CoreCruncher : Fast and Robust Construction of Core Genomes in Large Prokaryotic Data Sets

The core genome represents the set of genes shared by all, or nearly all, strains of a given population or species of prokaryotes. Inferring the core genome is integral to many genomic analyses, however, most methods rely on the comparison of all the pairs of genomes; a step that is becoming increasingly difficult given the massive accumulation of genomic data. Here, we present CoreCruncher; a program that robustly and rapidly constructs core genomes across hundreds or thousands of genomes. CoreCruncher does not compute all pairwise genome comparisons and uses a heuristic based on the distributions of identity scores to classify sequences as orthologs or paralogs/xenologs. Although it is much faster than current methods, our results indicate that our approach is more conservative than other tools and less sensitive to the presence of paralogs and xenologs. CoreCruncher is freely available from: https://github.com/lbobay/CoreCruncher. CoreCruncher is written in Python 3.7 and can also run on Python 2.7 without modification. It requires the python library Numpy and either Usearch or Blast. Certain options require the programs muscle or mafft.

59 BASIC BIOLOGICAL SCIENCES↗

Hybrid Storage Solution

With the rise of artificial intelligence and machine learning, data sets used to train models have become increasingly large. The availability, accessibility and integrity of large data sets has become important to the research conducted at Los Alamos National Laboratory. Ceph is a storage solution suitable for use with critical data because of its distributed nature and ability to keep multiple copies of a file in different locations. The amount of data means that bandwidth, latency, and cost are important factors and the reason most storage solutions are on-premises. However, there are distinct advantages to hosting services in the cloud, namely scalability and ease-of-use. In this paper, we explore the possibility of provisioning a hybrid Ceph cluster that leverages the benefits of both cloud architectures and on-premise performance.

97 MATHEMATICS AND COMPUTING↗

Ground State Energy Functional with Hartree–Fock Efficiency and Chemical Accuracy

We introduce the deep post Hartree–Fock (DeePHF) method, a machine learning-based scheme for constructing accurate and transferable models for the ground-state energy of electronic structure problems. DeePHF predicts the energy difference between results of highly accurate models such as the coupled cluster method and low accuracy models such as the Hartree–Fock (HF) method, using the ground-state electronic orbitals as the input. It preserves all the symmetries of the original high accuracy model. The added computational cost is less than that of the reference HF or DFT and scales linearly with respect to system size. We examine the performance of DeePHF on organic molecular systems using publicly available data sets and obtain the state-of-art performance, particularly on large data sets.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Data transfer for STAR grid jobs

The Solenoidal Tracker at RHIC (STAR) is a multipurpose experiment at the Relativistic Heavy Ion Collider (RHIC) with the primary goal to study the formation and properties of the quark-gluon plasma. STAR is an international collaboration of member institutions and laboratories from around the world. Yearly data-taking period produces PBytes of raw data collected by the experiment. STAR primarily uses its dedicated facility at BNL to process this data, but has routinely leveraged distributed systems, both high throughput (HTC) and high performance (HPC) computing clusters, to significantly augment the processing capacity available to the experiment. The ability to automate the efficient transfer of large data sets on reliable, scalable, and secure infrastructure is critical for any large-scale distributed processing campaign. For more than a decade, STAR computing has relied upon GridFTP with its x509-based authentication to build such data transfer systems and integrate them into its larger production workflow. The end of support by the community for both GridFTP and the x509 standard requires STAR to investigate other approaches to meet its distributed processing needs. In this study we investigate two multi-purpose data distribution systems, Globus.org and XRootD, as alternatives to GridFTP. We compare both their performance and the ease by which each service is integrated into the type of secure and automated data transfer systems STAR has previously built using GridFTP. The presented approach and study may be applicable to other distributed data processing use cases beyond STAR.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

SLIP (Surrogate Launching and Integration Platform)

SLIP (Surrogate Launching and Integration Platform) is a software ecosystem for downloading and running ML/AI benchmarks. SLIP automatically downloads and sets up the code and data running a benchmark. The large data sets and model code are cached locally with a specific ID within a designated cache directory.

Brown, Cade↗

From Images to Dark Matter: End-to-end Inference of Substructure from Hundreds of Strong Gravitational Lenses

Abstract Constraining the distribution of small-scale structure in our universe allows us to probe alternatives to the cold dark matter paradigm. Strong gravitational lensing offers a unique window into small dark matter halos (<10 10 M ⊙ ) because these halos impart a gravitational lensing signal even if they do not host luminous galaxies. We create large data sets of strong lensing images with realistic low-mass halos, Hubble Space Telescope (HST) observational effects, and galaxy light from HST’s COSMOS field. Using a simulation-based inference pipeline, we train a neural posterior estimator of the subhalo mass function (SHMF) and place constraints on populations of lenses generated using a separate set of galaxy sources. We find that by combining our network with a hierarchical inference framework, we can both reliably infer the SHMF across a variety of configurations and scale efficiently to populations with hundreds of lenses. By conducting precise inference on large and complex simulated data sets, our method lays a foundation for extracting dark matter constraints from the next generation of wide-field optical imaging surveys.

79 ASTRONOMY AND ASTROPHYSICS↗

Machine Learning Prediction of Tritium‐Helium Groundwater Ages in the Central Valley, California, USA

Abstract Groundwater ages provides insight into recharge rates, flow velocities, and vulnerability to contaminants. The ability to predict groundwater ages based on more accessible parameters via Machine Learning (ML) would advance our ability to guide sustainable management of groundwater resources. In this study, ML models were trained and tested on a large data set of tritium concentrations and tritium‐helium groundwater ages from the California Central Valley, a large groundwater basin with complex land use, irrigation, and water management practices. The ML models were trained on 63 features, including location, well construction information, landscape characteristics, and climate variables, water chemistry, and stable isotopes. The Bagging regressor method can accurately classify (F1‐score = 0.91) groundwater samples as either modern or pre‐modern whereas the accuracy of the ML prediction of continuous tritium‐helium groundwater ages is limited and explains only of the variability in this data set. In general, ML groundwater age prediction relies mostly on features related to (a) the source of groundwater recharge, (b) contaminant history, (c) aquifer materials, (d) well construction, and (e) geochemical reactions along flow paths.

54 ENVIRONMENTAL SCIENCES↗

Deeply learning deep inelastic scattering kinematics

We study the use of deep learning techniques to reconstruct the kinematics of the neutral current deep inelastic scattering (DIS) process in electron–proton collisions. In particular, we use simulated data from the ZEUS experiment at the HERA accelerator facility, and train deep neural networks to reconstruct the kinematic variables Q 2 and x. Our approach is based on the information used in the classical construction methods, the measurements of the scattered lepton, and the hadronic final state in the detector, but is enhanced through correlations and patterns revealed with the simulated data sets. We show that, with the appropriate selection of a training set, the neural networks sufficiently surpass all classical reconstruction methods on most of the kinematic range considered. Rapid access to large samples of simulated data and the ability of neural networks to effectively extract information from large data sets, both suggest that deep learning techniques to reconstruct DIS kinematics can serve as a rigorous method to combine and outperform the classical reconstruction methods.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Astronomaly at scale: searching for anomalies amongst 4 million galaxies

ABSTRACT Modern astronomical surveys are producing data sets of unprecedented size and richness, increasing the potential for high-impact scientific discovery. This possibility, coupled with the challenge of exploring a large number of sources, has led to the development of novel machine-learning-based anomaly detection approaches, such as astronomaly. For the first time, we test the scalability of astronomaly by applying it to almost 4 million images of galaxies from the Dark Energy Camera Legacy Survey. We use a trained deep learning algorithm to learn useful representations of the images and pass these to the anomaly detection algorithm isolation forest, coupled with astronomaly’s active learning method, to discover interesting sources. We find that data selection criteria have a significant impact on the trade-off between finding rare sources such as strong lenses and introducing artefacts into the data set. We demonstrate that active learning is required to identify the most interesting sources and reduce artefacts, while anomaly detection methods alone are insufficient. Using astronomaly, we find 1635 anomalies among the top 2000 sources in the data set after applying active learning, including eight strong gravitational lens candidates, 1609 galaxy merger candidates, and 18 previously unidentified sources exhibiting highly unusual morphology. Our results show that by leveraging the human–machine interface, astronomaly is able to rapidly identify sources of scientific interest even in large data sets.

Astronomy & Astrophysics↗

A variational encoder–decoder approach to precise spectroscopic age estimation for large Galactic surveys

Constraints on the formation and evolution of the Milky Way Galaxy require multidimensional measurements of kinematics, abundances, and ages for a large population of stars. Ages for luminous giants, which can be seen to large distances, are an essential component of studies of the Milky Way, but they are traditionally very difficult to estimate precisely for a large data set and often require careful analysis on a star-by-star basis in asteroseismology. Because spectra are easier to obtain for large samples, being able to determine precise ages from spectra allows for large age samples to be constructed, but spectroscopic ages are often imprecise and contaminated by abundance correlations. Here we present an application of a variational encoder–decoder on cross-domain astronomical data to solve these issues. The model is trained on pairs of observations from APOGEE and Kepler of the same star in order to reduce the dimensionality of the APOGEE spectra in a latent space while removing abundance information. The low dimensional latent representation of these spectra can then be trained to predict age with just ∼1000 precise seismic ages. We demonstrate that this model produces more precise spectroscopic ages (∼ 22 per cent overall, ∼ 11 per cent for red-clump stars) than previous data-driven spectroscopic ages while being less contaminated by abundance information (in particular, our ages do not depend on [α/M]). We create a public age catalogue for the APOGEE DR17 data set and use it to map the age distribution and the age-[Fe/H]-[α/M] distribution across the radial range of the Galactic disc.

79 ASTRONOMY AND ASTROPHYSICS↗