Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Sparse Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Quantum block encoding for one-pair semiseparable matrices

Quantum block encoding (QBE) is a crucial step in the development of most quantum algorithms, as it provides an embedding of a given matrix into a suitable larger unitary matrix. Historically, the development of efficient techniques for QBE has mostly focused on sparse matrices; less effort has been devoted to data-sparse (e.g., rank-structured) matrices. In this work we examine a particular case of rank structure, namely, one-pair semiseparable matrices. We present a new block encoding approach that relies on a suitable factorization of the given matrix as the product of triangular and diagonal factors. To encode the matrix, the algorithm needs $2\log(N)+7$ ancillary qubits. Assuming that the data input oracles can be implemented with polylogarithmic depth, or that a QRAM input model is available, our proposed method requires $\mathcal{O}({\rm polylog} (N))$ time and has an error of $\mathcal{O}(N^2)$, where $N$ is the matrix size.

Antonioli, Giacomo [Pisa U.; CERN] (ORCID:00090000↗

Results of a statistical approach to rainfall estimation using Nimbus 5 6.7 micrometers and 11.5 micrometers THIR data

Nimbus 5 6.7 mm and 11.5 mm temperature humidity infrared radiometer (THIR) data were used in a simple multiple regression scheme to test the feasibility of using these data to estimate hourly rainfall. Throughout the test area (85 W to 105 W and 45 N to 30 N) subareas (8 deg x 6 deg) were chosen from which point to point and areal statistics were obtained. Four subsets of data were used. The first consisted of only those surface stations indicating precipitation whose latitude and longitude coincided with the THIR grid points. A second used surface stations 0.1 degree from the THIR grid points. The third was a combination of subsets one and two. A reciprocal distance weighting scheme was used to derive precipitation values in data sparse areas. A fourth subset was made using these data combined with the data from subsets one and two. Point estimates resulted in negative correlations between estimated and grid derived "surface" precipitation. One degree areal estimates showed a slight improvement with a correlation coefficient of approximately 0.11. Single regression areal estimates resulted in correlations of approximately 0.11 and 0.20 for the 6.7 mm and 11.5 mm data respectively. These poor results were attributed to problems which are inherent in the satellite data (location errors, short temporal span of data, wavelength of sensors, etc.) and the lack of sufficient surface data to better verify the satellite estimate.

Ormsby, J. P.↗

Tropical cyclone track and genesis forecasting using satellite microwave sounder data

Although many dynamical and statistical prediction schemes are available to forecasters, tropical cyclone track errors are still large. One primary difficulty is that tropical cyclones exist over the data-sparse tropical oceans. Satellite sounders, however, routinely provide numerous data over these areas. Mean layer temperatures from the Scanning Microwave Spectrometer on board the Nimbus 6 satellite are decomposed using empirical orthogonal functions, and the expansion coefficients are related to deviations from the persistence forecast location, to speed change, to direction change and to intensity change. The significance of the regression equations is tested by a null hypothesis of zero correlation coefficient. It appears that significant information about tropical cyclone motion exists in the satellite-estimated mean layer temperatures, especially at upper levels. A physical interpretation of the statistical results is offered, and a one-storm-out independent test is used to test the stability of the equations. Finally, some further work is suggested.

Kidder, S. Q.↗

Integrating Applied Energy and BER Smart Data Capabilities to Develop a DOE Data Fabric for Energy-Water R&D

Focal Area(s): 1) Data acquisition and assimilation enabled by machine learning, AI, and advanced methods including experimental/network design/optimization, unsupervised learning (including deep learning), and hardware-related efforts involving AI (e.g., edge computing). Science Challenge: DOE R&D, including DOE’s Basic Energy Research (BER)’s Environmental Systems Science Division (EESSD) program and DOE’s applied energy research (AER) programs (EERE, FE, and NE) are producers and consumers of Earth systems datasets. This white paper focuses on the first topic area from the call in relation to how crosscutting resources and innovations from DOE’s EESSD and AER can be brought to bear to mutual benefit and more efficient energy-water, Earth system data resources through improved. The overarching challenge posed by this call focuses on how DOE can directly leverage artificial intelligence (AI) to engineer a substantial (paradigm-changing) improvement in Earth System Predictability? While stemming from DOE BER’s EESSD program, this is a challenge that is faced and also being addressed by DOE’s AER programs. Over the past decade plus, FE, EERE, and NE programs have made important strides towards addressing this need. These strides are in many ways highly complementary to EESSD’s MODEX efforts. Energy water systems spanning metocean to groundwater to surface water systems all are data driven whether for basic energy or applied energy. These are remote, multi-variate, complex natural, and in many cases engineered, systems. Key needs and challenges of both EESSD and AER include developing data-focused tools to enhance data search and discovery to fill in knowledge gaps (address sparse data challenge), and rapidly transform datasets, including disparate and multi-source data. Leveraging DOE on-premise computing (HPC, exascale) infrastructure supports the computing-intensive algorithms required to execute these data acquisition and transformation processes to derive enriched knowledge and data, driving AI/ML and big data analytics for these systems. The opportunity lies in combining BER and AER efforts to provide a more robust, advanced, efficient and complete computing data fabric to address energy-water data acquisition and assimilation needs which currently pose significant impediments to AI/ML predictions and research.

54 ENVIRONMENTAL SCIENCES↗

Determination of stratospheric temperature and height gradients from Nimbus 3 radiation data

To improve the specification of stratospheric, horizontal, temperature and geopotential height fields needed for high-flying aircraft, we derived a technique to estimate data between satellite tracks using interpolated infrared interferometer spectrometer 15-micron radiation data from Nimbus 3. The interpolation is based on the observed gradients of the medium resolution infrared 15-micron radiances between subsatellite tracks. The technique was verified with radiosonde data taken within 6 hr of the satellite data. Comparison indicates that the technique developed here produces analyses that are in general agreement with those from radiosonde data. In addition, this technique provides details over areas of sparse data not shown by conventional techniques.

Nicholas, G. W.↗

Location of Lunakhod 2 from laser ranging observations

Results of laser ranging attempts on Lunakhod 2 are presented. Laser ranging can only be accomplished at night, so only the straight-line distance between the end points of a daytime traverse can be determined. Uncertainties of positional determination are strongly dependent on the length of the tracking arc. The data, obtained in the first two lunar nights, are too sparse to yield reliable positional determinations. The results clearly illustrate the critical importance of a high-quality lunar ephemeris for the interpretation of sparse data.

Mulholland, J. D.↗

An initial approach to the utilization of VAS satellite sounding data

The primary objective is the development of methods to utilize satellite soundings. To achieve that, a more precise understanding of the characteristics of satellite derived soundings is required. The initial approach in the utilization of satellite soundings is to find a way to combine rawinsondes and satellite data into a unified data set. If this can be done in a meaningful way, satellite data could be integrated into existing data and interpretation would be enhanced. In this attempt, both rawinsonde and satellite soundings will be decomposed by the use of the Fourier cosine series. Then, each harmonic will be compared to get a better understanding of the representativeness of the satellite sounding data. Harmonics taken from satellite and rawinsonde soundings will be combined to provide a unified data set that can be used both over data sparse ocean areas as well as over land where rawinsonde data are available.

Shin, Kyung-Sup↗

Weather Intelligent Navigation Data and Models for Aviation Planning (WINDMAP)

WINDMAP addresses the emerging needs in the aviation community of providing real-time weather forecasting to improve the safety of low altitude aircraft operations. This is accomplished through the integration of real-time observations from autonomous systems, such as drones and urban air taxis, with numerical weather prediction models and flight management and safety systems. To solve this problem, several technical challenges have been identified. These include (1) developing autonomous UAS capable of conducting observations accurately and reliably; (2) determining the number and frequency of required observations and the sensitivity of these observations in data sparse regions of the lower atmosphere;(3) assimilating dense observational data into models in real-time with sufficient resolution and accuracy; (4) developing novel physics-based reduced order models capable of incorporating diverse data sets; and (5)integrating real-time forecasting into UTM and DAA (detect-and-avoid) architectures for path planning and navigation. The goal of this proposed effort is to demonstrate the value of using small UAS to collect measurements of the dynamic and thermodynamic properties of the lower atmosphere at scales that match or exceed the spatio-temporal resolution of today’s best numerical weather prediction models

Koushik Datta↗

Predicting fault slip via transfer learning

Abstract Data-driven machine-learning for predicting instantaneous and future fault-slip in laboratory experiments has recently progressed markedly, primarily due to large training data sets. In Earth however, earthquake interevent times range from 10’s-100’s of years and geophysical data typically exist for only a portion of an earthquake cycle. Sparse data presents a serious challenge to training machine learning models for predicting fault slip in Earth. Here we describe a transfer learning approach using numerical simulations to train a convolutional encoder-decoder that predicts fault-slip behavior in laboratory experiments. The model learns a mapping between acoustic emission and fault friction histories from numerical simulations, and generalizes to produce accurate predictions of laboratory fault friction. Notably, the predictions improve by further training the model latent space using only a portion of data from a single laboratory earthquake-cycle. The transfer learning results elucidate the potential of using models trained on numerical simulations and fine-tuned with small geophysical data sets for potential applications to faults in Earth.

58 GEOSCIENCES↗

A hill-sliding strategy for initialization of Gaussian clusters in the multidimensional space

A hill sliding technique was devised to extract Gaussian clusters from the multivariate probability density estimate of sample data for the first step of iterative unsupervised classification. Each cluster was assumed to posses a unimodal normal distribution. A clustering function proposed distinguished elements of a cluster under formation from the rest in the feature space. Initial clusters were extracted one by one according to the hill sliding tactics. A dimensionless cluster compactness parameter was proposed as a universal measure of cluster goodness and used satisfactorily in test runs with LANDSAT multispectral scanner data. The normalized divergence, defined by the cluster divergence divided by the entropy of the entire sample data, was utilized as a general separability measure between clusters. An overall clustering objective function was set forth in terms of cluster covariance matrices, from which the cluster compactness measure could be deduced. Minimal improvement of initial data partitioning was evaluated by this objective function in eliminating scattered sparse data points. The hill sliding clustering technique developed herein has the potential applicability to decomposition any multivariate mixture distribution into a number of unimodal distributions when an appropriate distribution function to the data set is employed.

Park, J. K.↗

Determining Greenland Ice Sheet Accumulation Rates from Radar Remote Sensing

An important component of NASA's Program for Arctic Regional Climate Assessment (PARCA) is a mass balance investigation of the Greenland Ice Sheet. The mass balance is calculated by taking the difference between the snow accumulation and the ice discharge of the ice sheet. Uncertainties in this calculation include the snow accumulation rate, which has traditionally been determined by interpolating data from ice core samples taken throughout the ice sheet. The sparse data associated with ice cores, coupled with the high spatial and temporal resolution provided by remote sensing, have motivated scientists to investigate relationships between accumulation rate and microwave observations.

Jezek, Kenneth C.↗

Machine-Learning of Nonlocal Kernels for Anomalous Subsurface Transport from Breakthrough Curves

Anomalous behavior is ubiquitous in subsurface solute transport due to the presence of high degrees of heterogeneity at different scales in the media. Although fractional models have been extensively used to describe the anomalous transport in various subsurface applications, their application is hindered by computational challenges. Simpler nonlocal models characterized by integrable kernels and finite interaction length represent a computationally feasible alternative to fractional models; yet, the informed choice of their kernel functions still remains an open problem. We propose a general data-driven framework for the discovery of optimal kernels on the basis of very small and sparse data sets in the context of anomalous subsurface transport. Using spatially sparse breakthrough curves recovered from fine-scale particle-density simulations, we learn the best coarse-scale nonlocal model using a nonlocal operator regression technique. Predictions of the breakthrough curves obtained using the optimal nonlocal model show good agreement with fine-scale simulation results even at locations and time intervals different from the ones used to train the kernel, confirming the excellent generalization properties of the proposed algorithm. A comparison with trained classical models and with black-box deep neural networks confirms the superiority of the predictive capability of the proposed model.

97 MATHEMATICS AND COMPUTING↗

Determination of stratospheric temperature and height gradients from nimbus 3 radiation data

To improve the specification of stratospheric horizontal temperature and geopotential height fields from satellite radiation data, needed for high flying aircraft, a technique was derived to estimate data between satellite tracks using interpolated IRIS 15-micron data from Nimbus III. The interpolation is based on the observed gradients of the MRIR 15-micron radiances between subsatellite tracks. The technique was verified with radiosonde data taken within 6 hours of the satellite data. The sample varied from 1126 pairs at low levels to 383 pairs at 10 mb using northern hemisphere data for June 15 to July 20, 1969. The data were separated into five latitude bands. The Rms temperature differences were generally from 2 to 5 C for all levels above 300 mb. From 500 to 300 mb RMS differences vary from 4 to 9C except at high latitudes which show values near 3C. The RMS differences between radiosonde heights and those calculated hydrostatically from the surface were from 30 to 280 meters increasing from the surface to 10 mb. Integration starting at 100 mb reduced the RMS difference in the stratosphere to 20 to 120 meters from 70 to 10 mb. From a comparison with actual operational maps at 50 and 10 mb, it appears the techniques developed produce analyses in general agreement with those from radiosonde data. In addition, they are able to indicate details over areas of sparse data not shown by conventional techniques.

Nicholas, G. W.↗

Trust-Enhancing Probabilistic Transfer Learning for Sparse and Noisy Data Environments

There is an increasing aspiration to utilize machine learning (ML) for various tasks of relevance to national security. ML models have thus far been mostly applied to tasks and domains that, while impactful, have sufficient volume of data. For predictive tasks of national security relevance, ML models of great capacity (ability to approximate nonlinear trends in input-output maps) are often needed to capture the complex underlying physics. However, scientific problems of relevance to national security are often accompanied by various sources of sparse and/or incomplete data, including experiments and simulations, across different regimes of operation, of varying degrees of fidelity, and include noise with different characteristics and/or intensity. State-of-the-art ML models, despite exhibiting superior performance on the task and domain they were trained on, may suffer detrimental loss in performance in such sparse data environments. This report summarizes the results of the Laboratory Directed Research and Development project entitled Trust-Enhancing Probabilistic Transfer Learning for Sparse and Noisy Data Environments. The objective of the project was to develop a new transfer learning (TL) framework that aims to adaptively blend the data across different sources in tackling one task of interest, resulting in enhanced trustworthiness of ML models for mission- and safety-critical systems. The proposed framework determines when it is worth applying TL and how much knowledge is to be transferred, despite uncontrollable uncertainties. The framework accomplishes this by leveraging concepts and techniques from the fields of Bayesian inverse modeling and uncertainty quantification, relying on strong mathematical foundations of probability and measure theories to devise new uncertainty-aware TL workflows.

97 MATHEMATICS AND COMPUTING↗

The importance of cycle-by-cycle data in performing rapid battery technology development and validation

Lithium-ion battery (LiB) technology is playing a crucial role in transforming the predominantly fossil fuel-based transportation and stationary storage sectors to achieve a low-carbon economy. Rapid innovation in the LiB materials to electrode to cell design is happening to satisfy the performance, life, and safety metrics required by those myriads of applications. Lately, advanced analytics, such as machine-learning or artificial intelligence (ML/AI) techniques, are being used more frequently to aid in expedited LiB technology development, performance validation, and life prediction. The success of these techniques often relies on a large volume of well-defined and high-quality battery test data. On the other hand, most battery developers and research and development (R&D) communities are still following a classical approach to develop batteries, which is running calendar- and/or cycle-aging tests, performing reference performance tests (RPTs), and conducting post-mortem analyses periodically without paying attention to the wealth of data often not collected during the calendar or cycle life aging tests. This sparse data collection approach is time- and resource-intensive, requiring data capture and evaluation of months to years of RPT data to diagnose accurate battery state of performance, health, and safety. Even so, the underlying aging modes and mechanisms can be missed. If collected properly, battery test data during cycling or calendaring can be efficiently combined with ML/AI techniques to create powerful tools in the rapid diagnosis of battery state of performance, health, and safety along with insights into underlying aging modes and mechanisms. In this report, we discuss the importance of effective cycle-by-cycle (CBC) data collection with example case studies. Within a reasonable timeframe, RPT data are often inadequate in capturing many of the crucial battery aging dynamics, which often predominantly show up in CBC test data. Finally, we also show examples of ML/AI techniques that use CBC data in rapid diagnosis and projection of LiB state of health (SOH) to motivate the scientific community in collecting and using CBC data to facilitate expeditious technology development and validation.

25 ENERGY STORAGE↗

Predictability of Malaria Transmission Intensity in the Mpumalanga Province, South Africa, Using Land Surface Climatology and Autoregressive Analysis

There has been increasing effort in recent years to employ satellite remotely sensed data to identify and map vector habitat and malaria transmission risk in data sparse environments. In the current investigation, available satellite and other land surface climatology data products are employed in short-term forecasting of infection rates in the Mpumalanga Province of South Africa, using a multivariate autoregressive approach. The climatology variables include precipitation, air temperature and other land surface states computed by the Off-line Land-Surface Global Assimilation System (OLGA) including soil moisture and surface evaporation. Satellite data products include the Normalized Difference Vegetation Index (NDVI) and other forcing data used in the Goddard Earth Observing System (GEOS-1) model. Predictions are compared to long- term monthly records of clinical and microscopic diagnoses. The approach addresses the high degree of short-term autocorrelation in the disease and weather time series. The resulting model is able to predict 11 of the 13 months that were classified as high risk during the validation period, indicating the utility of applying antecedent climatic variables to the prediction of malaria incidence for the Mpumalanga Province.

Grass, David↗

An algorithm for extraction of periodic signals from sparse, irregularly sampled data

Temporal gaps in discrete sampling sequences produce spurious Fourier components at the intermodulation frequencies of an oscillatory signal and the temporal gaps, thus significantly complicating spectral analysis of such sparsely sampled data. A new fast Fourier transform (FFT)-based algorithm has been developed, suitable for spectral analysis of sparsely sampled data with a relatively small number of oscillatory components buried in background noise. The algorithm's principal idea has its origin in the so-called 'clean' algorithm used to sharpen images of scenes corrupted by atmospheric and sensor aperture effects. It identifies as the signal's 'true' frequency that oscillatory component which, when passed through the same sampling sequence as the original data, produces a Fourier image that is the best match to the original Fourier space. The algorithm has generally met with succession trials with simulated data with a low signal-to-noise ratio, including those of a type similar to hourly residuals for Earth orientation parameters extracted from VLBI data. For eight oscillatory components in the diurnal and semidiurnal bands, all components with an amplitude-noise ratio greater than 0.2 were successfully extracted for all sequences and duty cycles (greater than 0.1) tested; the amplitude-noise ratios of the extracted signals were as low as 0.05 for high duty cycles and long sampling sequences. When, in addition to these high frequencies, strong low-frequency components are present in the data, the low-frequency components are generally eliminated first, by employing a version of the algorithm that searches for non-integer multiples of the discrete FET minimum frequency.

Wilcox, J. Z.↗

Stochastic modeling and statistical calibration with model error and scarce data

This paper introduces a procedure to assess the predictive accuracy of stochastic models subject to model error and sparse data. Model error is introduced as uncertainty on the coefficients of appropriate polynomial chaos expansions (PCE). The error associated with finite sample size allows us to conceive of these coefficients as statistics of the data that we describe as random variables whose influence on output quantities of interest is evaluated through the extended polynomial chaos expansion (EPCE). A Bayesian data assimilation scheme is introduced to update these expansions by considering the resulting nested chaos expansion as a hierarchical probabilistic model. Stochastic models of quantities of interest (QoI) are thus constructed and efficiently evaluated. Here, the Metropolis–Hastings Markov chain Monte Carlo procedure is used to sample the posterior. Two illustrative analytical and numerical problems are used to demonstrate the proposed approach.

Bayesian inference↗