Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Sparse Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Development of an MEMS ultrasonic microphone array system and its application to compressed wavefield imaging of concrete

Abstract Although contactless ultrasonic wavefield imaging shows potential for effective nondestructive inspection of various engineering materials, it has been rarely applied to concrete materials owing to technical challenges including low signal-to-noise ratio (SNR) caused by inherent heterogeneity of concrete. This paper presents development of a multi-channel MEMS ultrasonic microphone array system and its application to compressed wavefield imaging of concrete materials. The developed multi-channel MEMS ultrasonic microphone array system contains eight MEMS ultrasonic microphone elements and a signal conditioning circuit that enables measurements of ultrasonic signals with high SNR. A compressed sensing approach, based on the multiple measurement vector (MMV) concept, is applied to reconstruct a full dense ultrasonic wavefield data from sparsely sampled ultrasonic wavefield data. Experiments are carried out on a laboratory concrete sample to verify the performance of the developed MEMS microphone array system and proposed compressed sensing approach and then large-scale concrete samples to demonstrate practical application. The experimental results demonstrate that the developed MEMS microphone array system provides high-quality (SNR > 20 dB) ultrasonic data collected from concrete elements; furthermore, the proposed compressed sensing approach provides accurate reconstruction of dense wavefield data, as determined by peak signal-to-noise ratio (PSNR), from sparsely measured wavefield data with compression ratios up to 85% and PSNR above 25 dB in data collected form realistic large-scale concrete samples. By combining the MEMS array system and compressed sensing approach, the total ultrasonic data acquisition time needed to produce dense wavefield data can be significantly reduced.

Instruments & Instrumentation↗

Zero-truncated Poisson regression for sparse multiway count data corrupted by false zeros

Abstract We propose a novel statistical inference methodology for multiway count data that is corrupted by false zeros that are indistinguishable from true zero counts. Our approach consists of zero-truncating the Poisson distribution to neglect all zero values. This simple truncated approach dispenses with the need to distinguish between true and false zero counts and reduces the amount of data to be processed. Inference is accomplished via tensor completion that imposes low-rank tensor structure on the Poisson parameter space. Our main result shows that an $N$-way rank-$R$ parametric tensor $\boldsymbol{\mathscr{M}}\in (0,\infty )^{I\times \cdots \times I}$ generating Poisson observations can be accurately estimated by zero-truncated Poisson regression from approximately $IR^2\log _2^2(I)$ non-zero counts under the nonnegative canonical polyadic decomposition. Our result also quantifies the error made by zero-truncating the Poisson distribution when the parameter is uniformly bounded from below. Therefore, under a low-rank multiparameter model, we propose an implementable approach guaranteed to achieve accurate regression in under-determined scenarios with substantial corruption by false zeros. Several numerical experiments are presented to explore the theoretical results.

97 MATHEMATICS AND COMPUTING↗

Southern Ocean Seasonal Net Production from Satellite, Atmosphere, and Ocean Data Sets

A new climatology of monthly air-sea O2 flux was developed using the net air-sea heat flux as a template for spatial and temporal interpolation of sparse hydrographic data. The climatology improves upon the previous climatology of Najjar and Keeling in the Southern Hemisphere, where the heat-based approach helps to overcome limitations due to sparse data coverage. The climatology is used to make comparisons with productivity derived from CZCS images. The climatology is also used in support of an investigation of the plausible impact of recent global warming an oceanic O2 inventories.

Keeling, Ralph F.↗

Active Learning A Neural Network Model For Gold Clusters & Bulk From Sparse First Principles Training Data

Small metal clusters are of fundamental scientific interest and of tremendous significance in catalysis. These nanoscale clusters display diverse geometries and structural motifs depending on the cluster size; a knowledge of this size-dependent structural motifs and their dynamical evolution has been of longstanding interest. Given the high computational cost of first-principles calculations, molecular modeling and atomistic simulations such as molecular dynamics (MD) has proven to be an important complementary tool to aid this understanding. Classical MD typically employ predefined functional forms which limits their ability to capture such complex size-dependent structural and dynamical transformation. Neural Network (NN) based potentials represent flexible alternatives and in principle, well-trained NN potentials can provide high level of flexibility, transferability and accuracy on-par with the reference model used for training. A major challenge, however, is that NN models are interpolative and requires large quantities (similar to 10 4 or greater) of training data to ensure that the model adequately samples the energy landscape both near and far-from-equilibrium. A highly desirable goal is minimize the number of training data, especially if the underlying reference model is first-principles based and hence expensive. In this work, we introduce an active learning (AL) scheme that trains a NN model on-the-fly with minimal amount of first-principles based training data. Our AL workflow is initiated with a sparse training dataset (similar to 1 to 5 data points) and is updated on-the-fly via a Nested Ensemble Monte Carlo scheme that iteratively queries the energy landscape in regions of failure and updates the training pool to improve the network performance. Using a representative system of gold clusters, we demonstrate that our AL workflow can train a NN with similar to 500 total reference calculations. Using an extensive DFT test set of similar to 1100 configurations, we show that our AL-NN is able to accurately predict both the DFT energies and the forces for clusters of a myriad of different sizes. Our NN predictions are within 30 meV/atom and 40 meV/angstrom of the reference DFT calculations. Moreover, our AL-NN model also adequately captures the various size-dependent structural and dynamical properties of gold clusters in excellent agreement with DFT calculations and available experiments. We finally show that our AL-NN model also captures bulk properties reasonably well, even though they were not included in the training data.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

The Kepler DB, a Database Management System for Arrays, Sparse Arrays and Binary Data

The Kepler Science Operations Center stores pixel values on approximately six million pixels collected every 30-minutes, as well as data products that are generated as a result of running the Kepler science processing pipeline. The Kepler Database (Kepler DB) management system was created to act as the repository of this information. After one year of ight usage, Kepler DB is managing 3 TiB of data and is expected to grow to over 10 TiB over the course of the mission. Kepler DB is a non-relational, transactional database where data are represented as one dimensional arrays, sparse arrays or binary large objects. We will discuss Kepler DB's APIs, implementation, usage and deployment at the Kepler Science Operations Center.

McCauliff, Sean↗

Correlated Trajectory Uncertainty for Adaptive Sequential Decision Making

One of the great challenges with decision making tasks on real world systems is the fact that data is sparse and acquiring additional data is expensive. In these cases, it is often crucial to make a model of the environment to assist in making decisions. At the same time, limited data means that learned models are erroneous, making it just as important to equip the model with good predictive uncertainties. In the context of learning sequential decision making policies, these uncertainties can prove useful for informing which data to collect for the greatest improvement in policy performance \citep{mehta2021experimental, mehta2022exploration} or informing the policy about unsure regions of state and action space to avoid during test time \citep{yu2020mopo}. Additionally, assuming that realistic samples of the environment can be drawn, an adaptable policy can be trained that attempts to make optimal decisions for any given possible instance of the environment \citep{ghosh2022offline, chen2021offline}. In this work, we examine the so-called ``probabilistic neural network'' (PNN) model that is ubiquitous in model-based reinforcement learning (MBRL) works. We argue that while PNN models may have good marginal uncertainties, they form a distribution of non-smooth transition functions. Not only are these samples unrealistic and may hamper adaptability, but we also assert that this leads to poor uncertainty estimates when predicting multiple step trajectory estimates. To address this issue, we propose a simple sampling method that can be implemented on top of pre-existing models.We evaluate our sampling technique on a number of environments, including a realistic nuclear fusion task, and find that, not only do smooth transition function samples produce more calibrated uncertainties, but they also lead to better downstream performance for an adaptive policy.

Offline Reinforcement Learning↗

Investigation of serendipitious WFC sources

The serendipitious WFC sources under investigation, i.e., those which just happened to lie in the field of view while another object was being studied, were disappointing. The integration times were chosen to suit the primary target not the serendipitous targets. UX UMa, CZ Ori, BI Ori, WX Cet and AR And were not detected. A long (approximately 17 ksec) pointed observation of UX Uma (PI Wood) has since been carried out (February 1993) and the data is expected shortly. The other observations were much more successful. V471 Tau was observed with the WFC for 6.5 hrs with the S1 filter and 1.1 hrs with the S2b filter. It was easily detected with a count rate of 0.03 cps in the S1 filter and 0.15 cps in the S2b filter. The oscillations were seen, even before the data was folded, as was expected from preliminary results from the survey (Barstow et al. 1992). The pointed observations provide a much better phase coverage of the oscillations than did the survey data where the coverage was sparse. These data will be presented in a paper with the PSPC data (PI's Robinson and Shipman) and the pulse profiles in the different wavelengths will be compared.

Robinson, Edward L.↗

A Formal Messaging Notation for Alaskan Aviation Data

Data exchange is an increasingly important aspect of the National Airspace System. While many data communication channels have become more capable of sending and receiving data at higher throughput rates, there is still a need to use communication channels efficiently with limited throughput. The limitation can be based on technological issues, financial considerations, or both. This paper provides a complete description of several important aviation weather data in Abstract Syntax Notation format. By doing so, data providers can take advantage of Abstract Syntax Notation's ability to encode data in a highly compressed format. When data such as pilot weather reports, surface weather observations, and various weather predictions are compressed in such a manner, it allows for the efficient use of throughput-limited communication channels. This paper provides details on the Abstract Syntax Notation One (ASN.1) implementation for Alaskan aviation data, and demonstrates its use on real-world aviation weather data samples as Alaska has sparse terrestrial data infrastructure and data are often sent via relatively costly satellite channels.

communications↗

Application of VISSR Atmospheric Sounder (VAS) data in weather analysis

A technique which analyzes irregularly spaced satellite data is described. An experiment with rawinsonde and VISSR Atmospheric Sounder (VAS) radiance measurements collected on March 6-7, 1982 is conducted to reveal the applicability of the technique. The rawinsonde data are analyzed on a 16 x 12 grid using the two pass analysis scheme of Barnes (1973). A scheme similar to the Barnes (1973) procedure is employed to produce gridded analysis of VAS data over a 200 x 15000 km region in central part of the U.S. The use of a correction pass on the initial gridded field is described; the technique is extremely effective on uniformly spaced observations. The incorporation of the limited fine mesh model to the scheme to analyze data in sparse and cloudy regions is examined. A comparison of rawinsonde data with VAS data is provided. The technique proves effective for studying cloudy and sparse areas with VAS data and produces a four-dimensional data set with significant mesoscale structure.

Jedlovec, G. J.↗

Convolutional neural network based non-iterative reconstruction for accelerating neutron tomography *

Abstract Neutron computed tomography (NCT), a 3D non-destructive characterization technique, is carried out at nuclear reactor or spallation neutron source-based user facilities. Because neutrons are not severely attenuated by heavy elements and are sensitive to light elements like hydrogen, neutron radiography and computed tomography offer a complementary contrast to x-ray CT conducted at a synchrotron user facility. However, compared to synchrotron x-ray CT, the acquisition time for an NCT scan can be orders of magnitude higher due to lower source flux, low detector efficiency and the need to collect a large number of projection images for a high-quality reconstruction when using conventional algorithms. As a result of the long scan times for NCT, the number and type of experiments that can be conducted at a user facility is severely restricted. Recently, several deep convolutional neural network (DCNN) based algorithms have been introduced in the context of accelerating CT scans that can enable high quality reconstructions from sparse-view data. In this paper, we introduce DCNN algorithms to obtain high-quality reconstructions from sparse-view and low signal-to-noise ratio NCT data-sets thereby enabling accelerated scans. Our method is based on the supervised learning strategy of training a DCNN to map a low-quality reconstruction from sparse-view data to a higher quality reconstruction. Specifically, we evaluate the performance of two popular DCNN architectures—one based on using patches for training and the other on using the full images for training. We observe that both the DCNN architectures offer improvements in performance over classical multi-layer perceptron as well as conventional CT reconstruction algorithms. Our results illustrate that the DCNN can be a powerful tool to obtain high-quality NCT reconstructions from sparse-view data thereby enabling accelerated NCT scans for increasing user-facility throughput or enabling high-resolution time-resolved NCT scans.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Data traffic reduction schemes for sparse Cholesky factorizations

Load distribution schemes are presented which minimize the total data traffic in the Cholesky factorization of dense and sparse, symmetric, positive definite matrices on multiprocessor systems with local and shared memory. The total data traffic in factoring an n x n sparse, symmetric, positive definite matrix representing an n-vertex regular 2-D grid graph using n (sup alpha), alpha is equal to or less than 1, processors are shown to be O(n(sup 1 + alpha/2)). It is O(n(sup 3/2)), when n (sup alpha), alpha is equal to or greater than 1, processors are used. Under the conditions of uniform load distribution, these results are shown to be asymptotically optimal. The schemes allow efficient use of up to O(n) processors before the total data traffic reaches the maximum value of O(n(sup 3/2)). The partitioning employed within the scheme, allows a better utilization of the data accessed from shared memory than those of previously published methods.

Naik, Vijay K.↗

Pulse: An Outlier Sensitive Downsampling Algorithm For Timeseries Data

Pulse is a downsampling algorithm for timeseries data. Frequently datasets become so large that visualization tools and web browsers cannot effectively render graphics due to memory constraints. Downsampling algorithms are commonly applied to minimize the quantity of data required to visualize important features or trends in the data, but some datasets are composed by distinct enough features and trends that most existing downsampling algorithms fail to preserve them. Pule was developed to downsample timeseries data for galvanostatic stack test data at the Idaho National Laboratory. These datasets were composed by approximately 4 million records, most of them being extremely uniform. However, during relatively brief time periods when the stack test changes state, for example when the test article is powered on, or a load is added, the data produce sparse asymptotes. No existing downsampling algorithm was capable of preserving the sparse asymptotes in electrolysis stack test data. Instead, we develop a downsampling algorithm that preserves important outliers in data, and otherwise aggressively downsamples uniform data. The algorithm has applications in other domains like seismology, in the measurement of earthquakes, or astronomy, in the measurement of quasars or transit photometry.

Woodruff, Nathan [Idaho National Laboratory (INL),↗

An open-access simulated earthquake ground-motion database for an M7 Hayward Fault earthquake in the San Francisco Bay Region

Comprehensive understanding of earthquake ground motions, particularly in the near-fault region of large-magnitude events, is limited by gaps in strong-motion data. This challenge is prominent in areas with high seismic hazard but infrequent large earthquakes where data is sparse and difficult to interpret. These data limitations lead to uncertainties in the development of site-specific ground motions, which are crucial for engineering risk assessments. To address these challenges, physics-based regional-scale ground-motion simulations have been developed. With the emergence of exaflop-scale computing ecosystems, it is now possible to simulate regional earthquake processes at unprecedented fidelity and generate the large number of fault rupture realizations necessary to characterize both intra- and inter-event ground-motion variability. This article introduces a new database of simulated earthquake ground motions, created for applications in earthquake engineering, earthquake planning, and emergency response. The inaugural version of the database features simulated ground motions for a magnitude 7 Hayward Fault earthquake in the San Francisco Bay Region (SFBR), using the EarthQuake SIMulation (EQSIM) simulation framework and the Graves–Pitarka kinematic rupture model. The aim is to provide high-fidelity, spatially dense, three-component motions generated on the Department of Energy’s (DOE) newest generation of graphics processing unit (GPU)-accelerated supercomputers. These motions are being made openly available to the engineering, scientific, and disaster planning communities. In addition, this work develops protocols for the efficient dissemination of these large data sets and emphasizes community engagement to build confidence in their application. This article discusses the methodology behind the data, underlying software verification and validation, scalable data management, and a user interface for data access. The goal is to facilitate widespread use and elicit expert feedback to maximize the utility and exploitation of simulated motions. While the initial focus is on the San Francisco Region, simulations for additional regions will be added as the DOE program progresses.

Simulated ground-motion database↗