Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “large data sets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

PPINN: Parareal physics-informed neural network for time-dependent PDEs

Physics-informed neural networks (PINNs) encode physical conservation laws and prior physical knowledge into the neural networks, ensuring the correct physics is represented accurately while alleviating the need for supervised learning to a great degree. While effective for relatively short-term time integration, when long time integration of the time-dependent PDEs is sought, the time–space domain may become arbitrarily large and hence training of the neural network may become prohibitively expensive. To this end, we develop a parareal physics-informed neural network (PPINN), hence decomposing a long-time problem into many independent short-time problems supervised by an inexpensive/fast coarse-grained (CG) solver. In particular, the serial CG solver is designed to provide approximate predictions of the solution at discrete times, while initiate many fine PINNs simultaneously to correct the solution iteratively. There is a two-fold benefit from training PINNs with small-data sets rather than working on a large-data set directly, i.e., training of individual PINNs with small-data is much faster, while training the fine PINNs can be readily parallelized. Consequently, compared to the original PINN approach, the proposed PPINN approach may achieve a significant speed-up for long-time integration of PDEs, assuming that the CG solver is fast and can provide reasonable predictions of the solution, hence aiding the PPINN solution to converge in just a few iterations. To investigate the PPINN performance on solving time-dependent PDEs, we first apply the PPINN to solve the Burgers equation, and subsequently we apply the PPINN to solve a two-dimensional nonlinear diffusion–reaction equation. Furthermore, our results demonstrate that PPINNs converge in a few iterations with significant speed-ups proportional to the number of time-subdomains employed.

42 ENGINEERING↗

Building Large-Scale U.S. Synthetic Electric Distribution System Models

Rapid increases in distributed energy resources on distribution systems are prompting research efforts to improve and evaluate electric power distribution algorithms; however, there is a shortage of realistic, large-scale, U.S.-style test systems for the evaluation of such advanced algorithms. Some available tools to build large-scale test systems are of European style, and their application to studies focused on the United States might not be desirable given very different characteristics between the European and U.S. distribution designs. Motivated by this need, this paper develops detailed algorithms to build large-scale U.S. distribution systems and incorporates them in a new Reference Network Model, RNM-US. The approach starts with information from street maps and a catalog with electric equipment that includes power lines, transformers, voltage regulators, capacitors, and switching devices. The paper presents the algorithms through an illustrative case study of the different steps that comprise the process of building a synthetic distribution grid. Finally, the paper presents a medium- and a large-scale data set covering 10 million electrical nodes and 120,000 km of power lines, demonstrating the applicability of the proposed method to build very large-scale synthetic distribution systems.

27 ARPA - Advanced Research Projects Agency-Energy↗

Spatio-Temporal Surrogates for Interaction of a Jet with High Explosives: Part I - Analysis with a Small Sample Size

Computer simulations, especially of complex phenomena, can be expensive, requiring high-performance computing resources. Often, to understand a phenomenon, multiple simulations are run, each with a different set of simulation input parameters. These data are then used to create an interpolant, or surrogate, relating the simulation outputs to the corresponding inputs. When the inputs and outputs are scalars, a simple machine learning model can suffice. However, when the simulation outputs are vector valued, available at locations in two or three spatial dimensions, often with a temporal component, creating a surrogate is more challenging. In this report, we use a two-dimensional problem of a jet interacting with high explosives to understand how we can build high-quality surrogates. The characteristics of our data set are unique - the vector-valued outputs from each simulation are available at over two million spatial locations; each simulation is run for a relatively small number of time steps; the size of the computational domain varies with each simulation; and resource constraints limit the number of simulations we can run. We show how we analyze these extremely large data-sets, set the parameters for the algorithms used in the analysis, and use simple ways to improve the accuracy of the spatio-temporal surrogates without substantially increasing the number of simulations required.

97 MATHEMATICS AND COMPUTING↗

Metabolic Pathway Engineering

Modern microbial and enzyme engineering and their advancement are increasingly dependent on the marriage of a wide range of sophisticated technologies. For students entering the field of biotechnology, the outlook is indeed daunting. Expertise at levels beyond that of simple familiarity will be needed to conduct competitive research. It goes without saying that to be competitive, all new researchers in this field will need basic preparation in molecular biology, biochemistry, and genetics. Further, experience working with concepts and experimental tools in enzyme biochemistry and kinetics, gene editing, computational metabolic pathway modeling, experimental pathway flux analysis, and computational clustering tools to process complex data sets will be vital for success. We speculate that most biotechnology researchers in early career at this time will build teams of collaborators to address these disparate science fields rather than attempt to become experts in one lab.

BASIC BIOLOGICAL SCIENCES,BIOMASS FUELS↗

Improved Parametric Models for Explosion Pressure Signals Derived From Large Datasets

Accurate recording and characterization of explosion-induced pressure signals are key components of the forensic analysis of explosion events in the atmosphere. Parametric overpressure models based on several key waveform features (peak overpressure, positive pulse duration, and impulse) are widely used to estimate explosion energy in terms of trinitrotoluene equivalent yield. However, those models are often developed by a limited dataset, including only a few events or recordings at relatively short propagation distances. Here, we develop empirical waveform-parameter models based on a regression analysis of a large set of data curated from four chemical explosion experiments including 16 detonations. We measured peak overpressure and impulse for positive and negative phases from 1000 pressure signals recorded at local ranges (<20 km) with scaled distance up to 8000 m/kg 1/3 . Additionally, the measured waveform parameters showed large variation with respect to observing distances indicating the effects of atmospheric propagation. In this study, a second-order polynomial model was used in a least-squares regression to account for those propagation effects and to improve data fitting. In addition to model parameters for waveform features, we also determined range-dependent model uncertainties based on data variance. The model uncertainty represents the prediction error of our models and can be critical to evaluating the uncertainty of yield estimate.

58 GEOSCIENCES↗

Cosmological constraints from higher redshift gamma-ray burst, H ii starburst galaxy, and quasar (and other) data

ABSTRACT We use higher redshift gamma-ray burst (GRB), H ii starburst galaxy (H iiG), and quasar angular size (QSO-AS) measurements to constrain six spatially flat and non-flat cosmological models. These three sets of cosmological constraints are mutually consistent. Cosmological constraints from a joint analysis of these data sets are largely consistent with currently accelerating cosmological expansion and with cosmological constraints derived from a combined analysis of Hubble parameter (H(z)) and baryon acoustic oscillation (BAO, with Planck-determined baryonic matter density) measurements. A joint analysis of the H(z) + BAO + QSO-AS + H iiG + GRB data provides fairly model-independent determinations of the non-relativistic matter density parameter $\Omega _{\rm m_0}=0.313\pm 0.013$ and the Hubble constant $H_0=69.3\pm 1.2\, \rm {km \, s^{-1} \, Mpc^{-1}}$. These data are consistent with the dark energy being a cosmological constant and with spatial hypersurfaces being flat, but they do not rule out mild dark energy dynamics or a little spatial curvature. We also investigate the effect of including quasar flux measurements in the mix and find no novel conclusions.

79 ASTRONOMY AND ASTROPHYSICS↗

Accelerating Structure–Property Relationship Discovery with Multimodal Machine Learning and Self-Driving Microscopy

Microscopy combined with local spectroscopy is widely used to correlate nanoscale structure with functional properties in materials, but conventional measurements rely heavily on human-selected sampling locations and predefined targets, limiting data set diversity and the potential for discovery. Here, we present a framework that integrates autonomous microscopy with dual-novelty deep kernel learning (DN-DKL) for adaptive data acquisition and a dual variational autoencoder (VAE) for representation learning. DN-DKL actively guides the microscopy toward structurally and spectroscopically novel regions, enabling efficient collection of large spectral data sets. Dual-VAE embeds local structures and spectroscopic responses into a shared latent manifold that serves as a structure–property relationship map. We applied this framework for the investigation of halide perovskite films by using conductive atomic force microscopy. The results reveal distinct hysteresis behaviors that are linked to specific nanoscale structural motifs, including grain boundary junction points that show hysteresis under different bias conditions and asymmetric grain boundaries that suppress the charge transport. This framework establishes a general strategy that leverages the complementary strengths of self-driving microscopy, machine learning, and human expertise to accelerate scientific discovery in functional materials.

atomic force microscopy↗

Computationally Efficient Learning of Large Scale Dynamical Systems: A Koopman Theoretic Approach

In recent years there has been a considerable drive towards data-driven analysis, discovery and control of dynamical systems. To this end, operator theoretic methods, namely, Koopman operator methods have gained a lot of interest. In general, the Koopman operator is obtained as a solution to a least-squares problem, and as such, the Koopman operator can be expressed as a closed-form solution that involves the computation of Moore-Penrose inverse of a matrix. For high dimensional systems and also if the size of the obtained data-set is large, the computation of the Moore-Penrose inverse becomes computationally challenging. In this paper, we provide an algorithm for computing the Koopman operator for high dimensional systems in a time-efficient manner. We further demonstrate the efficacy of the proposed approach on two different systems, namely a network of coupled oscillators (with state-space dimension up to 2500) and IEEE 68 bus system (with state-space dimension 204 and up to 24,000 time-points).

Sinha, Subhrajit↗

A deep learning approach for semantic segmentation of unbalanced data in electron tomography of catalytic materials

In computed TEM tomography, image segmentation represents one of the most basic tasks with implications not only for 3D volume visualization, but more importantly for quantitative 3D analysis. In case of large and complex 3D data sets, segmentation can be an extremely difficult and laborious task, and thus has been one of the biggest hurdles for comprehensive 3D analysis. Heterogeneous catalysts have complex surface and bulk structures, and often sparse distribution of catalytic particles with relatively poor intrinsic contrast, which possess a unique challenge for image segmentation, including the current state-of-the-art deep learning methods. To tackle this problem, we apply a deep learning-based approach for the multi-class semantic segmentation of a γ-Alumina/Pt catalytic material in a class imbalance situation. Specifically, we used the weighted focal loss as a loss function and attached it to the U-Net’s fully convolutional network architecture. We assessed the accuracy of our results using Dice similarity coefficient (DSC), recall, precision, and Hausdorff distance (HD) metrics on the overlap between the ground-truth and predicted segmentations. Our adopted U-Net model with the weighted focal loss function achieved an average DSC score of 0.96 ± 0.003 in the γ-Alumina support material and 0.84 ± 0.03 in the Pt NPs segmentation tasks. We report an average boundary-overlap error of less than 2 nm at the 90th percentile of HD for γ-Alumina and Pt NPs segmentations. The complex surface morphology of γ-Alumina and its relation to the Pt NPs were visualized in 3D by the deep learning-assisted automatic segmentation of a large data set of high-angle annular dark-field (HAADF) scanning transmission electron microscopy (STEM) tomography reconstructions.

36 MATERIALS SCIENCE↗

Evaluating skill in predicting the Interdecadal Pacific Oscillation in initialized decadal climate prediction hindcasts in E3SMv1 and CESM1 using two different initialization methods and a small set of start years

Abstract It is a daunting challenge to conduct initialized hindcasts with enough ensemble members and associated start years to form a drifted climatology from which to compute the anomalies necessary to quantify the skill of the hindcasts when compared to observations. This limits the ability to experiment with case studies and other applications where only a few initial years are needed. Here we run a set of hindcasts with CESM1 and E3SMv1 using two different initialization methods for a limited set of start years and use the respective uninitialized free-running historical simulations to form the model climatologies. Since the drifts from the observed initial states in the hindcasts toward the uninitialized model state are large and rapid, after a few years the drifted initialized models approach the uninitialized model climatological errors. Therefore, hindcasts from the limited start years can use the uninitialized climatology to represent the drifted model states after about lead year 3, providing a means to compute forecast anomalies in the absence of a large hindcast sample. There is comparable skill for predicting spatial patterns of multi-year Pacific sea surface temperature anomalies in the domain of the Interdecadal Pacific Oscillation using this method compared to the conventional methodology with a large hindcast data set, though there is a model dependence to the drifts in the two initialization methods.

54 ENVIRONMENTAL SCIENCES↗

Carbonate clumped isotope analysis (Δ 47 ) of 21 carbonate standards determined via gas‐source isotope‐ratio mass spectrometry on four instrumental configurations using carbonate‐based standardization and multiyear data sets

Rationale Clumped isotope geochemistry examines the pairing or clumping of heavy isotopes in molecules and provides information about the thermodynamic and kinetic controls on their formation. The first clumped isotope measurements of carbonate minerals were first published 15 years ago, and since then, interlaboratory offsets have been observed, and laboratory and community practices for measurement, data analysis, and instrumentation have evolved. Here we briefly review historical and recent developments for measurements, share Tripati Lab practices for four different instrument configurations, test a recently published proposal for carbonate‐based standardization on multiple instruments using multi‐year data sets, and report values for 21 different carbonate standards that allow for recalculations of previously published data sets. Methods We examine data from 4628 standard measurements on Thermo MAT 253 and Nu Perspective IS mass spectrometers, using a common acid bath (90°C) and small‐sample (70°C) individual reaction vessels. Each configuration was investigated by treating some standards as anchors (working standards) and the remainder as unknowns (consistency standards). Results We show that different acid digestion systems and mass spectrometer models yield indistinguishable results when instrument drift is well characterized. For linearity correction, mixed gas‐and‐carbonate standardization or carbonate‐only standardization yields similar results. No difference is observed in the use of three or eight working standards for the construction of transfer functions. Conclusions We show that all configurations yield similar results if instrument drift is robustly characterized and validate a recent proposal for carbonate‐based standardization using large multiyear data sets. Δ 47 values are reported for 21 carbonate standards on both the absolute reference frame (ARF; also refered to as the Carbon Dioxide Equilibrated Scale or CDES) and the new InterCarb‐Carbon Dioxide Equilibrium Scale (I‐CDES) reference frame, facilitating intercomparison of data from a diversity of labs and instrument configurations and restandardization of a broad range of sample sets between 2006, when the first carbonate measurements were published, and the present.

58 GEOSCIENCES↗

Baryon acoustic oscillations in thin redshift shells from BOSS DR12 and eBOSS DR16 galaxies

ABSTRACT In an age of large astronomical data sets and severe cosmological tensions, the case for model independent analyses is compelling. We present a set of 14 baryon acoustic oscillations measurements in thin redshift shells with $3\,\mathrm{ per} \,\mathrm{ cent}$ precision that were obtained by analysing BOSS DR12 and eBOSS DR16 galaxies in the redshift range 0.32 < z < 0.66. Thanks to the use of thin shells, the analysis is carried out using just redshifts and angles so that the fiducial model is only introduced when considering the mock catalogues, necessary for the covariance matrix estimation and the pipeline validation. We compare our measurements, with and without supernova data, to the corresponding constraints from Planck 2018, finding good compatibility. A Monte Python module for this likelihood is available at github.com/ranier137/angularBAO.

79 ASTRONOMY AND ASTROPHYSICS↗

Practical galaxy morphology tools from deep supervised representation learning

Astronomers have typically set out to solve supervised machine learning problems by creating their own representations from scratch. We show that deep learning models trained to answer every Galaxy Zoo DECaLS question learn meaningful semantic representations of galaxies that are useful for new tasks on which the models were never trained. We exploit these representations to outperform several recent approaches at practical tasks crucial for investigating large galaxy samples. The first task is identifying galaxies of similar morphology to a query galaxy. Given a single galaxy assigned a free text tag by humans (e.g. ‘#diffuse’), we can find galaxies matching that tag for most tags. The second task is identifying the most interesting anomalies to a particular researcher. Our approach is 100 per cent accurate at identifying the most interesting 100 anomalies (as judged by Galaxy Zoo 2 volunteers). The third task is adapting a model to solve a new task using only a small number of newly labelled galaxies. Models fine-tuned from our representation are better able to identify ring galaxies than models fine-tuned from terrestrial images (ImageNet) or trained from scratch. We solve each task with very few new labels; either one (for the similarity search) or several hundred (for anomaly detection or fine-tuning). This challenges the longstanding view that deep supervised methods require new large labelled data sets for practical use in astronomy. To help the community benefit from our pretrained models, we release our fine-tuning code zoobot. Zoobot is accessible to researchers with no prior experience in deep learning.

79 ASTRONOMY AND ASTROPHYSICS↗

L-PBF High-Throughput Data Pipeline Approach for Multi-modal Integration

Abstract Metal-based additive manufacturing requires active monitoring solutions for assessing part quality. Multiple sensors and data streams, however, generate large heterogeneous data sets that are impractical for manual assessment and characterization. In this work, an automated pipeline is developed that enables feature extraction from high-speed camera video and multi-modal data analysis. The framework removes the need for manual assessment through the utilization of deep learning techniques and training models in a weakly supervised paradigm. We demonstrate this pipeline’s capability over 700,000 high-speed camera frames. The pipeline successfully extracts melt pool and spatter geometries and links them to corresponding pyrometry, radiography, and processparameter information. 715 individual prints are examined to reveal melt pool areas that exceeds 0.07 mm 2 and pyrometry signal over a threshold (375 pyrometry units) were more likely to have defects. These automated processes enable massive throughput of characterization techniques.

36 MATERIALS SCIENCE↗

Audacity of huge: overcoming challenges of data scarcity and data quality for machine learning in computational materials discovery

Machine learning (ML)-accelerated discovery requires large amounts of high-fidelity data to reveal predictive structure–property relationships. For many properties of interest in materials discovery, the challenging nature and high cost of data generation has resulted in a data landscape that is both scarcely populated and of dubious quality. Data-driven techniques starting to overcome these limitations include the use of consensus across functionals in density functional theory, the development of new functionals or accelerated electronic structure theories, and the detection of where computationally demanding methods are most necessary. When properties cannot be reliably simulated, large experimental data sets can be used to train ML models. In the absence of manual curation, increasingly sophisticated natural language processing and automated image analysis are making it possible to learn structure–property relationships from the literature. Finally, models trained on these data sets will improve as they incorporate community feedback.

36 MATERIALS SCIENCE↗

Classification of animal sounds in a hyperdiverse rainforest using convolutional neural networks with data augmentation

To protect tropical forest biodiversity, we need to be able to detect it reliably, cheaply, and at scale. Automated detection of sound producing animals from passively recorded soundscapes via machine-learning approaches is a promising technique towards this goal, but it is constrained by the necessity of large training data sets. Using soundscapes from a tropical forest in Borneo and a Convolutional Neural Network model (CNN), we investigate i) the minimum viable training data set size for accurate prediction of call types (‘sonotypes’), and ii) the extent to which data augmentation and transfer learning can overcome the issue of small and imbalanced training data sets. We found that even relatively high sample sizes (>80 per sonotype) lead to mediocre accuracy, which however improved significantly with data augmentation and transfer learning, including at extremely small sample sizes (3 per sonotype), regardless of taxonomic group or call characteristics. Neither transfer learning nor data augmentation alone achieved high accuracy. Our results suggest that transfer learning and data augmentation could make the use of CNNs to classify species’ vocalizations feasible even for small soundscape-based projects with many rare species. Retraining our open-source model requires only basic programming skills which makes it possible for individual conservation initiatives to match their local context, in order to enable more evidence-informed management of biodiversity.

54 ENVIRONMENTAL SCIENCES↗

Source term estimation using noble gas and aerosol samples

Algorithms that estimate the location, time, and magnitude of a point-source atmospheric release using remotely sampled air concentrations typically use data for a single chemical or radioactive isotope. Here, a Bayesian algorithm is presented that uses data from multiple radioactive isotopes that are all released in the same short-duration event. Data from noble gas and aerosol samplers can be used simultaneously in the model. Application to a large synthetic data set using four isotopes shows the new algorithm generally gives more accurate location and time estimates than a comparable model using a single isotope.

54 ENVIRONMENTAL SCIENCES↗

Ligand-Based Compound Activity Prediction via Few-Shot Learning

Predicting the activities of new compounds against biophysical or phenotypic assays based on the known activities of one or a few existing compounds is a common goal in early stage drug discovery. This problem can be cast as a “few-shot learning” challenge, and prior studies have developed few-shot learning methods to classify compounds as active versus inactive. However, the ability to go beyond classification and rank compounds by expected affinity is more valuable. We describe Few-Shot Compound Activity Prediction (FS-CAP), a novel neural architecture trained on a large bioactivity data set to predict compound activities against an assay outside the training set, based on only the activities of a few known compounds against the same assay. Our model aggregates encodings generated from the known compounds and their activities to capture assay information and uses a separate encoder for the new compound whose activity is to be predicted. The new method provides encouraging results relative to traditional chemical-similarity-based techniques as well as other state-of-the-art few-shot learning methods in tests on a variety of ligand-based drug discovery settings and data sets.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗