Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Sparse Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Leveraging Prior Concept Learning Improves Generalization From Few Examples in Computational Models of Human Object Recognition

Humans quickly and accurately learn new visual concepts from sparse data, sometimes just a single example. The impressive performance of artificial neural networks which hierarchically pool afferents across scales and positions suggests that the hierarchical organization of the human visual system is critical to its accuracy. These approaches, however, require magnitudes of order more examples than human learners. We used a benchmark deep learning model to show that the hierarchy can also be leveraged to vastly improve the speed of learning. We specifically show how previously learned but broadly tuned conceptual representations can be used to learn visual concepts from as few as two positive examples; reusing visual representations from earlier in the visual hierarchy, as in prior approaches, requires significantly more examples to perform comparably. These results suggest techniques for learning even more efficiently and provide a biologically plausible way to learn new visual concepts from few examples.

Rule, Joshua S.↗

Large spatiotemporal variability in aerosol properties over central Argentina during the CACTI field campaign

Abstract. Few field campaigns with extensive aerosol measurements have been conducted over continental areas in the Southern Hemisphere. To address this data gap and better understand the interactions of convective clouds and the surrounding environment, extensive in situ and remote sensing measurements were collected during the Cloud, Aerosol, and Complex Terrain Interactions (CACTI) field campaign conducted between October 2018 and April 2019 over the Sierras de Córdoba range of central Argentina. This study describes measurements of aerosol number, size, composition, mixing state, and cloud condensation nuclei (CCN) collected on the ground and from a research aircraft during 7 weeks of the campaign. Large spatial and multiday variations in aerosol number, size, composition, and CCN were observed due to transport from upwind sources controlled by mesoscale to synoptic-scale meteorological conditions. Large vertical wind shears, back trajectories, single-particle measurements, and chemical transport model predictions indicate that different types of emissions and source regions, including biogenic emissions and biomass burning from the Amazon and anthropogenic emissions from Chile and eastern Argentina, contribute to aerosols observed during CACTI. Repeated aircraft measurements near the boundary layer top reveal strong spatial and temporal variations in CCN and demonstrate that understanding the complex co-variability of aerosol properties and clouds is critical to quantify the impact of aerosol–cloud interactions. In addition to quantifying aerosol properties in this data-sparse region, these measurements will be valuable to evaluate predictions over the midlatitudes of South America and improve parameterized aerosol processes in local, regional, and global models.

54 ENVIRONMENTAL SCIENCES↗

Discrete-Direct Model Calibration and Uncertainty Propagation Method Confirmed on Multi-Parameter Plasticity Model Calibrated to Sparse Random Field Data

A discrete direct (DD) model calibration and uncertainty propagation approach is explained and demonstrated on a 4-parameter Johnson-Cook (J-C) strain-rate dependent material strength model for an aluminum alloy. The methodology's performance is characterized in many trials involving four random realizations of strain-rate dependent material-test data curves per trial, drawn from a large synthetic population. The J-C model is calibrated to particular combinations of the data curves to obtain calibration parameter sets which are then propagated to “Can Crush” structural model predictions to produce samples of predicted response variability. These are processed with appropriate sparse-sample uncertainty quantification (UQ) methods to estimate various statistics of response with an appropriate level of conservatism. This is tested on 16 output quantities (von Mises stresses and equivalent plastic strains) and it is shown that important statistics of the true variabilities of the 16 quantities are bounded with a high success rate that is reasonably predictable and controllable. The DD approach has several advantages over other calibration-UQ approaches like Bayesian inference for capturing and utilizing the information obtained from typically small numbers of replicate experiments in model calibration situations—especially when sparse replicate functional data are involved like force–displacement curves from material tests. The DD methodology is straightforward and efficient for calibration and propagation problems involving aleatory and epistemic uncertainties in calibration experiments, models, and procedures.

42 ENGINEERING↗

Modeling of Hidden Structures Using Sparse Chemical Shift Data from NMR Relaxation Dispersion

NMR relaxation dispersion measurements report on conformational changes occurring on the μs-ms timescale. Chemical shift information derived from relaxation dispersion can be used to generate structural models of weakly populated alternative conformational states. Current methods to obtain such models rely on determining the signs of chemical shift changes between the conformational states, which are difficult to obtain in many situations. Here, we use a “sample and select” method to generate relevant structural models of alternative conformations of the C-terminal-associated region of Escherichia coli dihydrofolate reductase (DHFR), using only unsigned chemical shift changes for backbone amides and carbonyls ( 1 H, 15 N, and 13 C'). We find that CS-Rosetta sampling with unsigned chemical shift changes generates a diversity of structures that are sufficient to characterize a minor conformational state of the C-terminal region of DHFR. The excited state differs from the ground state by a change in secondary structure, consistent with previous predictions from chemical shift hypersurfaces and validated by the x-ray structure of a partially humanized mutant of E. coli DHFR (N23PP/G51PEKN). The results demonstrate that the combination of fragment modeling with sparse chemical shift data can determine the structure of an alternative conformation of DHFR sampled on the μs-ms timescale. Such methods will be useful for characterizing alternative states, which can potentially be used for in silico drug screening, as well as contributing to understanding the role of minor states in biology and molecular evolution.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Efficient Data Compression for 3D Sparse TPC via Bicephalous Convolutional Autoencoder

Real-time data collection and analysis in large experimental facilities present a great challenge across multiple domains, including high energy physics, nuclear physics, and cosmology. To address this, machine learning (ML)-based methods for real-time data compression have drawn significant attention. However, unlike natural image data, such as CIFAR and ImageNet that are relatively small-sized and continuous, scientific data often come in as three-dimensional 3D data volumes at high rates with high sparsity (many zeros) and non-Gaussian value distribution. This makes direct application of popular ML compression methods, as well as conventional data compression methods, suboptimal. To address these obstacles, this work introduces a dual-head autoencoder to resolve sparsity and regression simultaneously, called Bicephalous Convolutional AutoEncoder (BCAE). This method shows advantages both in compression fidelity and ratio compared to traditional data compression methods, such as MGARD, SZ, and ZFP. To achieve similar fidelity, the best performer among the traditional methods can reach only half the compression ratio of BCAE. Moreover, a thorough ablation study of the BCAE method shows that a dedicated segmentation decoder improves the reconstruction.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Development of an MEMS ultrasonic microphone array system and its application to compressed wavefield imaging of concrete

Abstract Although contactless ultrasonic wavefield imaging shows potential for effective nondestructive inspection of various engineering materials, it has been rarely applied to concrete materials owing to technical challenges including low signal-to-noise ratio (SNR) caused by inherent heterogeneity of concrete. This paper presents development of a multi-channel MEMS ultrasonic microphone array system and its application to compressed wavefield imaging of concrete materials. The developed multi-channel MEMS ultrasonic microphone array system contains eight MEMS ultrasonic microphone elements and a signal conditioning circuit that enables measurements of ultrasonic signals with high SNR. A compressed sensing approach, based on the multiple measurement vector (MMV) concept, is applied to reconstruct a full dense ultrasonic wavefield data from sparsely sampled ultrasonic wavefield data. Experiments are carried out on a laboratory concrete sample to verify the performance of the developed MEMS microphone array system and proposed compressed sensing approach and then large-scale concrete samples to demonstrate practical application. The experimental results demonstrate that the developed MEMS microphone array system provides high-quality (SNR > 20 dB) ultrasonic data collected from concrete elements; furthermore, the proposed compressed sensing approach provides accurate reconstruction of dense wavefield data, as determined by peak signal-to-noise ratio (PSNR), from sparsely measured wavefield data with compression ratios up to 85% and PSNR above 25 dB in data collected form realistic large-scale concrete samples. By combining the MEMS array system and compressed sensing approach, the total ultrasonic data acquisition time needed to produce dense wavefield data can be significantly reduced.

Instruments & Instrumentation↗

Zero-truncated Poisson regression for sparse multiway count data corrupted by false zeros

Abstract We propose a novel statistical inference methodology for multiway count data that is corrupted by false zeros that are indistinguishable from true zero counts. Our approach consists of zero-truncating the Poisson distribution to neglect all zero values. This simple truncated approach dispenses with the need to distinguish between true and false zero counts and reduces the amount of data to be processed. Inference is accomplished via tensor completion that imposes low-rank tensor structure on the Poisson parameter space. Our main result shows that an $N$-way rank-$R$ parametric tensor $\boldsymbol{\mathscr{M}}\in (0,\infty )^{I\times \cdots \times I}$ generating Poisson observations can be accurately estimated by zero-truncated Poisson regression from approximately $IR^2\log _2^2(I)$ non-zero counts under the nonnegative canonical polyadic decomposition. Our result also quantifies the error made by zero-truncating the Poisson distribution when the parameter is uniformly bounded from below. Therefore, under a low-rank multiparameter model, we propose an implementable approach guaranteed to achieve accurate regression in under-determined scenarios with substantial corruption by false zeros. Several numerical experiments are presented to explore the theoretical results.

97 MATHEMATICS AND COMPUTING↗

Active Learning A Neural Network Model For Gold Clusters & Bulk From Sparse First Principles Training Data

Small metal clusters are of fundamental scientific interest and of tremendous significance in catalysis. These nanoscale clusters display diverse geometries and structural motifs depending on the cluster size; a knowledge of this size-dependent structural motifs and their dynamical evolution has been of longstanding interest. Given the high computational cost of first-principles calculations, molecular modeling and atomistic simulations such as molecular dynamics (MD) has proven to be an important complementary tool to aid this understanding. Classical MD typically employ predefined functional forms which limits their ability to capture such complex size-dependent structural and dynamical transformation. Neural Network (NN) based potentials represent flexible alternatives and in principle, well-trained NN potentials can provide high level of flexibility, transferability and accuracy on-par with the reference model used for training. A major challenge, however, is that NN models are interpolative and requires large quantities (similar to 10 4 or greater) of training data to ensure that the model adequately samples the energy landscape both near and far-from-equilibrium. A highly desirable goal is minimize the number of training data, especially if the underlying reference model is first-principles based and hence expensive. In this work, we introduce an active learning (AL) scheme that trains a NN model on-the-fly with minimal amount of first-principles based training data. Our AL workflow is initiated with a sparse training dataset (similar to 1 to 5 data points) and is updated on-the-fly via a Nested Ensemble Monte Carlo scheme that iteratively queries the energy landscape in regions of failure and updates the training pool to improve the network performance. Using a representative system of gold clusters, we demonstrate that our AL workflow can train a NN with similar to 500 total reference calculations. Using an extensive DFT test set of similar to 1100 configurations, we show that our AL-NN is able to accurately predict both the DFT energies and the forces for clusters of a myriad of different sizes. Our NN predictions are within 30 meV/atom and 40 meV/angstrom of the reference DFT calculations. Moreover, our AL-NN model also adequately captures the various size-dependent structural and dynamical properties of gold clusters in excellent agreement with DFT calculations and available experiments. We finally show that our AL-NN model also captures bulk properties reasonably well, even though they were not included in the training data.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Correlated Trajectory Uncertainty for Adaptive Sequential Decision Making

One of the great challenges with decision making tasks on real world systems is the fact that data is sparse and acquiring additional data is expensive. In these cases, it is often crucial to make a model of the environment to assist in making decisions. At the same time, limited data means that learned models are erroneous, making it just as important to equip the model with good predictive uncertainties. In the context of learning sequential decision making policies, these uncertainties can prove useful for informing which data to collect for the greatest improvement in policy performance \citep{mehta2021experimental, mehta2022exploration} or informing the policy about unsure regions of state and action space to avoid during test time \citep{yu2020mopo}. Additionally, assuming that realistic samples of the environment can be drawn, an adaptable policy can be trained that attempts to make optimal decisions for any given possible instance of the environment \citep{ghosh2022offline, chen2021offline}. In this work, we examine the so-called ``probabilistic neural network'' (PNN) model that is ubiquitous in model-based reinforcement learning (MBRL) works. We argue that while PNN models may have good marginal uncertainties, they form a distribution of non-smooth transition functions. Not only are these samples unrealistic and may hamper adaptability, but we also assert that this leads to poor uncertainty estimates when predicting multiple step trajectory estimates. To address this issue, we propose a simple sampling method that can be implemented on top of pre-existing models.We evaluate our sampling technique on a number of environments, including a realistic nuclear fusion task, and find that, not only do smooth transition function samples produce more calibrated uncertainties, but they also lead to better downstream performance for an adaptive policy.

Offline Reinforcement Learning↗

Convolutional neural network based non-iterative reconstruction for accelerating neutron tomography *

Abstract Neutron computed tomography (NCT), a 3D non-destructive characterization technique, is carried out at nuclear reactor or spallation neutron source-based user facilities. Because neutrons are not severely attenuated by heavy elements and are sensitive to light elements like hydrogen, neutron radiography and computed tomography offer a complementary contrast to x-ray CT conducted at a synchrotron user facility. However, compared to synchrotron x-ray CT, the acquisition time for an NCT scan can be orders of magnitude higher due to lower source flux, low detector efficiency and the need to collect a large number of projection images for a high-quality reconstruction when using conventional algorithms. As a result of the long scan times for NCT, the number and type of experiments that can be conducted at a user facility is severely restricted. Recently, several deep convolutional neural network (DCNN) based algorithms have been introduced in the context of accelerating CT scans that can enable high quality reconstructions from sparse-view data. In this paper, we introduce DCNN algorithms to obtain high-quality reconstructions from sparse-view and low signal-to-noise ratio NCT data-sets thereby enabling accelerated scans. Our method is based on the supervised learning strategy of training a DCNN to map a low-quality reconstruction from sparse-view data to a higher quality reconstruction. Specifically, we evaluate the performance of two popular DCNN architectures—one based on using patches for training and the other on using the full images for training. We observe that both the DCNN architectures offer improvements in performance over classical multi-layer perceptron as well as conventional CT reconstruction algorithms. Our results illustrate that the DCNN can be a powerful tool to obtain high-quality NCT reconstructions from sparse-view data thereby enabling accelerated NCT scans for increasing user-facility throughput or enabling high-resolution time-resolved NCT scans.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Pulse: An Outlier Sensitive Downsampling Algorithm For Timeseries Data

Pulse is a downsampling algorithm for timeseries data. Frequently datasets become so large that visualization tools and web browsers cannot effectively render graphics due to memory constraints. Downsampling algorithms are commonly applied to minimize the quantity of data required to visualize important features or trends in the data, but some datasets are composed by distinct enough features and trends that most existing downsampling algorithms fail to preserve them. Pule was developed to downsample timeseries data for galvanostatic stack test data at the Idaho National Laboratory. These datasets were composed by approximately 4 million records, most of them being extremely uniform. However, during relatively brief time periods when the stack test changes state, for example when the test article is powered on, or a load is added, the data produce sparse asymptotes. No existing downsampling algorithm was capable of preserving the sparse asymptotes in electrolysis stack test data. Instead, we develop a downsampling algorithm that preserves important outliers in data, and otherwise aggressively downsamples uniform data. The algorithm has applications in other domains like seismology, in the measurement of earthquakes, or astronomy, in the measurement of quasars or transit photometry.

Woodruff, Nathan [Idaho National Laboratory (INL),↗

An open-access simulated earthquake ground-motion database for an M7 Hayward Fault earthquake in the San Francisco Bay Region

Comprehensive understanding of earthquake ground motions, particularly in the near-fault region of large-magnitude events, is limited by gaps in strong-motion data. This challenge is prominent in areas with high seismic hazard but infrequent large earthquakes where data is sparse and difficult to interpret. These data limitations lead to uncertainties in the development of site-specific ground motions, which are crucial for engineering risk assessments. To address these challenges, physics-based regional-scale ground-motion simulations have been developed. With the emergence of exaflop-scale computing ecosystems, it is now possible to simulate regional earthquake processes at unprecedented fidelity and generate the large number of fault rupture realizations necessary to characterize both intra- and inter-event ground-motion variability. This article introduces a new database of simulated earthquake ground motions, created for applications in earthquake engineering, earthquake planning, and emergency response. The inaugural version of the database features simulated ground motions for a magnitude 7 Hayward Fault earthquake in the San Francisco Bay Region (SFBR), using the EarthQuake SIMulation (EQSIM) simulation framework and the Graves–Pitarka kinematic rupture model. The aim is to provide high-fidelity, spatially dense, three-component motions generated on the Department of Energy’s (DOE) newest generation of graphics processing unit (GPU)-accelerated supercomputers. These motions are being made openly available to the engineering, scientific, and disaster planning communities. In addition, this work develops protocols for the efficient dissemination of these large data sets and emphasizes community engagement to build confidence in their application. This article discusses the methodology behind the data, underlying software verification and validation, scalable data management, and a user interface for data access. The goal is to facilitate widespread use and elicit expert feedback to maximize the utility and exploitation of simulated motions. While the initial focus is on the San Francisco Region, simulations for additional regions will be added as the DOE program progresses.

Simulated ground-motion database↗

Inferring demographic and selective histories from population genomic data using a 2-step approach in species with coding-sparse genomes: an application to human data

Abstract The demographic history of a population, and the distribution of fitness effects (DFE) of newly arising mutations in functional genomic regions, are fundamental factors dictating both genetic variation and evolutionary trajectories. Although both demographic and DFE inference has been performed extensively in humans, these approaches have generally either been limited to simple demographic models involving a single population, or, where a complex population history has been inferred, without accounting for the potentially confounding effects of selection at linked sites. Taking advantage of the coding-sparse nature of the genome, we propose a 2-step approach in which coalescent simulations are first used to infer a complex multi-population demographic model, utilizing large non-functional regions that are likely free from the effects of background selection. We then use forward-in-time simulations to perform DFE inference in functional regions, conditional on the complex demography inferred and utilizing expected background selection effects in the estimation procedure. Throughout, recombination and mutation rate maps were used to account for the underlying empirical rate heterogeneity across the human genome. Importantly, within this framework it is possible to utilize and fit multiple aspects of the data, and this inference scheme represents a generalized approach for such large-scale inference in species with coding-sparse genomes.

Soni, Vivak (ORCID:0000000294969562)↗

Thermodynamic Consistent Neural Networks for Learning Material Interfacial Mechanics

For multilayer materials in thin substrate systems, interfacial failure is one of the most challenges. The traction-separation relations (TSR) quantitatively describe the mechanical behavior of a material interface undergoing openings, which is critical to understand and predict interfacial failures under complex loadings. However, existing theoretical models have limitations on enough complexity and flexibility to well learn the real-world TSR from experimental observations. A neural network can fit well along with the loading paths but often fails to obey the laws of physics, due to a lack of experimental data and understanding of the hidden physical mechanism. In this paper, we propose a thermodynamic consistent neural network (TCNN) approach to build a data-driven model of the TSR with sparse experimental data. The TCNN leverages recent advances in physics-informed neural networks (PINN) that encode prior physical information into the loss function and efficiently train the neural networks using automatic differentiation. We investigate three thermodynamic consistent principles, i.e., positive energy dissipation, steepest energy dissipation gradient, and energy conservative loading path. All of them are mathematically formulated and embedded into a neural network model with a novel defined loss function. A real-world experiment demonstrates the superior performance of TCNN, and we find that TCNN provides an accurate prediction of the whole TSR surface and significantly reduces the violated prediction against the laws of physics.

Zhang, Jiaxin↗