Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Sparse Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Transitioning NPOESS Data to Weather Offices: The SPoRT Paradigm with EOS Data

Real-time satellite information provides one of many data sources used by NWS weather forecast offices (WFOs) to diagnose current weather conditions and to assist in short-term forecast preparation. While GOES satellite data provides relatively coarse spatial resolution coverage of the continental U.S. on a 10-15 minute repeat cycle, polar orbiting imagery has the potential to provide snapshots of weather conditions at high-resolution in many spectral channels. Additionally, polar orbiting sounding data can provide additional information on the thermodynamic structure of the atmosphere in data sparse regions of at asynoptic observation times. The NASA Short-term Prediction Research and Transition (SPoRT) project has demonstrated the utility of polar orbiting MODIS and AIRS data on the Terra and Aqua satellites to improve weather diagnostics and short-term forecasting on the regional and local scales. SPoRT scientists work directly forecasters at selected WFOS in the Southern Region (SR) to help them ingest these unique data streams into their AWIPS system, understand how to use the data (through on-site and distance learn techniques), and demonstrate the utility of these products to address significant forecast problems. This process also prepares forecasters for the use of similar observational capabilities from NPOESS operational sensors. NPOESS environmental data records (EDRs) from the Visible 1 Infrared Imager I Radiometer Suite (VIIRS), the Cross-track Infrared Sounder (CrlS) and Advanced Technology Microwave Sounder (ATMS) instruments and additional value-added products produced by NESDIS will be available in near real-time and made available to WFOs to extend their use of NASA EOS data into the NPOESS era. These new data streams will be integrated into the NWs's new AWIPS II decision support tools. The AWIPS I1 system to be unveiled in WFOs in 2009 will be a JAVA-based decision support system which preserves the functionality of the existing systems and offers unique development opportunities for new data sources and applications in the Service Orientated Architecture ISOA) environment. This paper will highlight some of the SPoRT activities leading to the integration of VllRS and CrIS/ATMS data into the display capabilities of these new systems to support short-term forecasting problems at WFOs.

Jedlovec, Gary↗

On the sensitivity of numerical weather prediction to remotely sensed marine surface wind data - A simulation study

The reported investigation has the objective to assess the potential impact on numerical weather prediction (NWP) of remotely sensed surface wind data. Other investigations conducted with similar objectives have not been satisfactory in connection with a use of procedures providing an unrealistic distribution of initial errors. In the current study, care has been taken to duplicate the actual distribution of information in the conventional observing system, thus shifting the emphasis from accuracy of the data to the data coverage. It is pointed out that this is an important consideration in assessing satellite observing systems since experience with sounder data has shown that improvements in forecasts due to satellite-derived information is due less to a general error reduction than to the ability to fill data-sparse regions. The reported study concentrates on the evaluation of the observing system simulation experimental design and on the assessment of the potential of remotely sensed marine surface wind data.

Cane, M. A.↗

Compatibility study of the Magsat data and aeromagnetic data in the eastern Piedmont of US

Data from a 2 day period recorded by Magsat were used to produce world magnetic maps of the scalar total field and three vector component total fields. Subtracting the reference field of Magsat 6/80, a scalar anomalous field and three vector component anomalous fields were also mapped. After removing 718 bad points from the original data, every fifth point was picked for contouring. While the main geomagnetic field of the Earth is surprisingly well mapped considering the short data period, the anomaly maps suffer from data sparseness. The entire Magsat file collected at altitudes of 500-700 m in nonmountainous terrain and at 900-1,000 m in mountainous terrain was averaged to reduce the total data to 6,500 measurements, yielding a 0.1 deg sampling interval along the flight path. A U.S. aeromagnetic anomaly surface map was produced and the field was upward continued to a 300 km altitude. Differences in anomaly structure between the POGO data and the map produced were attributed to insufficient removal of the reference field. Reprocessing of the data using the GSFC reference field (9/80-2) should remove the low harmonic field and improve the anomalous field structure.

Won, I. J.↗

TRMM Observations of Lightning and Rainfall

A multi-sensor algorithm is proposed that uses total lightning observations in conjunction,with conventional weather satellite imagery to develop proportionality relationships that can be used to improve space-time estimates of rainfall in data sparse regions. Previous studies have examined the relationships between rainfall and cloud-to-ground lightning only. The proposed algorithm is developed from relationships developed between total lightning and rainfall data collected at the TRMM ground validation site at Kennedy Space Center, Florida and elsewhere. The algorithm is evaluated throughout the tropics with data collected by the TRMM Lightning Imaging Sensor (LIS) and the other TRMM instruments. Based on earlier studies of relationships among total lightning, passive microwave ice scattering signatures, and cloud top height, this algorithm is expected to improve rainfall estimates from geosynchronous orbit. A lightning sensor is currently being designed for a future flight on the GOES satellite.

Goodman, S. J.↗

Generation of Data-Driven Expected Energy Models for Photovoltaic Systems

Although unique expected energy models can be generated for a given photovoltaic (PV) site, a standardized model is also needed to facilitate performance comparisons across fleets. Current standardized expected energy models for PV work well with sparse data, but they have demonstrated significant over-estimations, which impacts accurate diagnoses of field operations and maintenance issues. This research addresses this issue by using machine learning to develop a data-driven expected energy model that can more accurately generate inferences for energy production of PV systems. Irradiance and system capacity information was used from 172 sites across the United States to train a series of models using Lasso linear regression. The trained models generally perform better than the commonly used expected energy model from international standard (IEC 61724-1), with the two highest performing models ranging in model complexity from a third-order polynomial with 10 parameters (Radj2 = 0.994) to a simpler, second-order polynomial with 4 parameters (Radj2=0.993), the latter of which is subject to further evaluation. Subsequently, the trained models provide a more robust basis for identifying potential energy anomalies for operations and maintenance activities as well as informing planning-related financial assessments. We conclude with directions for future research, such as using splines to improve model continuity and better capture systems with low (≤1000 kW DC) capacity.

14 SOLAR ENERGY↗

Event Definition for the Automated Detection of Nuclear Proliferation Activities

In FY2020, Savannah River National Laboratory (SRNL) in collaboration with the Discovery Analytics Center (DAC) at Virginia Polytechnic Institute and State University (VT) began developing a demonstration prototype system that uses multiple machine learning and data analytic methods on largescale open data sources to identify new, developing, or undeclared nuclear programs. One of the most challenging aspects of applying machine learning techniques to such a problem is the high likelihood of extremely sparse data from disparate sources. To overcome this challenge, the current work will use a strategic combination of supervised, semi-supervised, and unsupervised learning techniques to ingest and fuse data streams to make a forecast of nuclear activities in a targeted geospatial location. Identifying potential data sources and training supervised learning algorithms is dependent upon the development of a robust foundation of targeted event domains that fundamentally define the nuclear activities of interest. This report documents the definition of a hierarchical structure for both nuclear activity and event domains that will be used to guide the research team in development or use of existing semantic dictionaries that are instrumental to searching, parsing, and categorizing events for the forecasting system’s use.

97 MATHEMATICS AND COMPUTING↗

Model Choice Metrics to Optimize Profile-QSAR Performance

Predicting molecular activity against protein targets is difficult because of the paucity of experimental data. Approaches like multitask modeling and collaborative filtering seek to improve model accuracy by leveraging results from multiple targets, but are limited because different compounds are measured with different assays, leading to sparse data matrices. Profile-QSAR (pQSAR) 2.0 addresses this problem by fitting a series of partial least squares models for each target, using as features the predictions from single-task models on the remaining targets. Here, this method has been shown to produce better results than single task and multitask models. However, the factors determining the success of pQSAR 2.0 have as yet not been characterized. In this paper we examine the experimental conditions that lead to better pQSAR models. We limit the amount of data available to the method by retraining with decreasing amounts of data and explore the model’s ability to generalize to compounds that have never been assayed. Finally, we look at the properties of training data needed to demonstrate pQSAR improvement.

Biological and medical sciences, Computer science↗

The Effect of the Saharan Air Layer on the Formation of Hurricane Isabel (2003) Simulated with AIRS Data

The crucial physics of how the atmosphere really accomplishes the tropical cyclogenesis process is still poorly understood. The presence of the Saharan Air Layer (SAL), an elevated mixed layer of warm and dry air that extends from Africa to the tropical Atlantic and contains a substantial amount of mineral dust, adds more complexity to the tropical cyclogenesis process in the Atlantic basin. The impact of the SAL on tropical cyclogenesis is still uncertain. Karyampudi and Carlson (1988) conclude that a strong SAL can potentially aid tropical cyclone development while Dunion and Velden (2004) argue that the SAL generally inhibits tropical cyclogenesis and intensification. Advancing our understanding of the physical mechanisms of tropical cyclogenesis and the associated roles of the SAL strongly depends on the improvement in the observations over the data-sparse ocean areas. After the Atmospheric Infrared Sounder (AIRS), the Advanced Microwave Sounding Unit (AMSU), and the microwave Humidity Sounder of Brazil (HSB) were launched with the NASA Aqua satellite in 2002, new data products retrieved from the AIRS suite became available for studying the effect of the warm, dry air mass associated with the SAL (referred to as the thermodynamic effect). The vertical profiles of the AIRS retrieved temperature and humidity provide an unprecedented opportunity to examine the thermodynamic effect of the SAL. The observational data can be analyzed and assimilated into numerical models, in which the model thermodynamic state is allowed to relax to the observed state from AIRS data. The objective of this study is to numerically demonstrate that the thermodynamic effect of the SAL on the formation of Hurricane Isabel (2003) can be largely simulated through nudging of the AIRS data.

Wu, iguang↗

Liftoff and Transition Database Generation for Launch Vehicles Using Data-Fusion-Based Modeling

A data fusion technique for merging multiple data sources with differing fidelity and resolution was developed to support the production of aerodynamic line load databases for the Liftoff and Transition (LOT) flight phase of the Space Launch System (SLS). The technique uses a reduced order model based on a high-fidelity line load data set from Computational Fluid Dynamics (CFD) to predict solutions for a much larger solution space. Even higher-fidelity force and moment information (from wind-tunnel tests) is then used to adjust the model. The adjustment uses constrained optimization through the method of Lagrange multipliers in order to minimize the deviation of the line load distribution from the spatially-dense CFD solution, while ensuring that the integrated force and moment values match those observed in physical wind tunnel measurements. Though the wind-tunnel data are operationally-dense (available at many flow conditions), they are spatially coarse (as only the overall forces and moments are available). Conversely, CFD for such complex configurations is expensive, and thus operationally sparse. Data fusion techniques are necessary to make the most efficient use of available information, delivering accurate results within time and resource constraints.

Wignall, T. J.↗

The inverse gravimetric problem in gravity modelling

One of the main purposes of geodesy is to determine the gravity field of the Earth in the space outside its physical surface. This purpose can be pursued without any particular knowledge of the internal density even if the exact shape of the physical surface of the Earth is not known, though this seems to entangle the two domains, as it was in the old Stoke's theory before the appearance of Molodensky's approach. Nevertheless, even when large, dense and homogeneous data sets are available, it was always recognized that subtracting from the gravity field the effect of the outer layer of the masses (topographic effect) yields a much smoother field. This is obviously more important when a sparse data set is bad so that any smoothing of the gravity field helps in interpolating between the data without raising the modeling error, this approach is generally followed because it has become very cheap in terms of computing time since the appearance of spectral techniques. The mathematical description of the Inverse Gravimetric Problem (IGP) is dominated mainly by two principles, which in loose terms can be formulated as follows: the knowledge of the external gravity field determines mainly the lateral variations of the density; and the deeper the density anomaly giving rise to a gravity anomaly, the more improperly posed is the problem of recovering the former from the latter. The statistical relation between rho and n (and its inverse) is also investigated in its general form, proving that degree cross-covariances have to be introduced to describe the behavior of rho. The problem of the simultaneous estimate of a spherical anomalous potential and of the external, topographic masses is addressed criticizing the choice of the mixed collection approach.

Sanso, F.↗

Reply to Comment by Peterie Et Al. on “Accelerated Fill‐Up of the Arbuckle Group Aquifer and Links to U.S. Midcontinent Seismicity”

Abstract Peterie et al. question one observation in our paper: associating pressure increases to injection volumes at distances of up to 25 km from an injection well. In this reply, we show that the comment misunderstands our analysis and the evidence that led to this conclusion. We also show that gauge‐depth‐corrected pressures, used by the authors to produce statewide pressure maps, are discrepant with the static fluid level data, provided in our original compilation and analysis. The discrepancies are a result of the pressure correction method employed, which naïvely substitutes formation pressure for bottomhole pressure to calculate wellbore fluid density. Their linearly interpolated pressure maps, based on sparse data, contain interpolation and extrapolation artifacts that contradict injection trends in the state, the Theis solution, and the superposition principle. We reiterate that pressure and static fluid level increases in Class I wells existed prior to 2013, most notably in central Kansas, where recent earthquakes are cited in the comment as evidence of a pressure plume emanating from the Kansas‐Oklahoma border, 90 km away. We show that the space‐time pattern of seismicity in this area is inconsistent with a northward propagating pressure plume and, instead, seismicity appears to be centered on and near a cluster of high‐rate injection wells, two of which are among the highest rate wells in the state. These observations, along with recent M4 + earthquakes during continued decreases in wastewater injection in southern Kansas and northern Oklahoma, question the usefulness of the comment for understanding and managing societally significant earthquakes.

Ansari, Esmail↗

NASA to launch NOAA's GOES-C earth monitoring satellite

NASA's launch of the GOES-C geostationary satellite from Kennedy Space Center, Florida is planned for June 16, 1978. The launch vehicle is a three stage Delta 2914. As its contribution, GOES-C will contribute information from a data sparse area of the world centered in the Indian Ocean. GOES-C will replace GOES-1 and will become GOES-3 once it has successfully orbited at 35,750 kilometers (22,300 miles). NASA's Spaceflight Tracking and Data Network (STDN) will provide support for the mission. Included in the article are: (1) Delta launch vehicle statistics, first, second and third stages; (2) Delta/GOES-C major launch events; (3) Launch operations; (4) Delta/GOES-C personnel.

Source record↗

Solving Inverse Stochastic Problems from Discrete Particle Observations Using the Fokker--Planck Equation and Physics-Informed Neural Networks

The Fokker--Planck (FP) equation governing the evolution of the probability density function (PDF) is applicable to many disciplines, but it requires specification of the coefficients for each case, which can be functions of space-time and not just constants and hence require the development of a data-driven modeling approach. When the data available is directly on the PDF, there exist methods for inverse problems that can be employed to infer the coefficients and thus determine the FP equation and subsequently obtain its solution. Herein, we address a more realistic scenario, where only sparse data are given on the particles' positions at a few time instants, which are not sufficient to accurately construct directly the PDF even at those times from existing methods, e.g., kernel estimation algorithms. To this end, we develop a general framework based on physics-informed neural networks (PINNs) that introduces a new loss function using the Kullback--Leibler divergence to connect the stochastic samples with the FP equation to simultaneously learn the equation and infer the multidimensional PDF at all times. In particular, we consider two types of inverse problems, type I, where the FP equation is known but the initial PDF is unknown, and type II, in which, in addition to the unknown initial PDF, the drift and diffusion terms are also unknown. In both cases, we investigate problems with either Brownian or Lévy noise or a combination of both. Here, we demonstrate the new PINN framework in detail in the one-dimensional (1D) case, but we also provide results for up to five dimensions demonstrating that we can infer both the FP equation and dynamics simultaneously at all times with high accuracy using only very few discrete observations of the particles.

97 MATHEMATICS AND COMPUTING↗

Machine Learning Inference of Random Medium Properties

Earth materials are heterogeneous across a range of spatial scales, but the resolvability of small structures is limited by sparse data coverage, noise, bandlimitedness, and other difficulties. In practice, heterogeneities below a certain size cannot be recovered from seismic data except through statistical medium descriptions, which even then can be difficult to uniquely determine. To improve the characterization of such heterogeneities, we develop a novel supervised machine learning (ML) model that provides insight about the recoverability of statistical medium properties from elastic waveform data and succeeds despite cycle-skipping and other challenges well known from elastic waveform inversion. We demonstrate the approach using random media generated by superimposing self-affine random variations on homogeneous and layered background structures. After training on sparsely-recorded, high-frequency waveforms from hundreds of different random medium realizations, we show the ability of our ML model to recover correlation lengths and other statistical properties of interest to near-surface and crustal seismology, among other fields. For frequency passbands and spatial offsets encountered in seismology, Gaussian correlation lengths and the amplitude of the random variations relative to the background model are recovered even in challenging scenarios involving unknown medium parameters, complex crustal structures, and low signal-to-noise ratio. In comparison, von Kármán correlation lengths, which are related to larger-wavelength variations of the medium than Gaussian correlation lengths, are not as well recovered. These results provide one of the first and most systematic investigations of the recoverability of statistical properties of heterogeneities below the resolution limit of deterministic seismic tomography, and suggest practical ML strategies for high-frequency waveform seismology.

58 GEOSCIENCES↗

Water Vapor Winds and Their Application to Climate Change Studies

The retrieval of satellite-derived winds and moisture from geostationary water vapor imagery has matured to the point where it may be applied to better understanding longer term climate changes that were previously not possible using conventional measurements or model analysis in data-sparse regions. In this paper, upper-tropospheric circulation features and moisture transport covering ENSO periods are presented and discussed. Precursors and other detectable interannual climate change signals are analyzed and compared to model diagnosed features. Estimates of winds and humidity over data-rich regions are used to show the robustness of the data and its value over regions that have previously eluded measurement.

Jedlovec, Gary J.↗

Satellite-Derived Water Vapor Winds for Regional Climate Studies

The retrieval of winds and humidity in the upper-troposphere has matured to the point where it may now be possible to better understand and diagnose regional climate variations from geostationary satellites than from conventional measurements or model analysis, especially in data sparse regions. In this poster paper, upper-tropospheric circulation features and moisture transport covering ENSO periods are presented and discussed. Precursor and other detectable interannual climate signals are analyzed and compared to model diagnosed features. Estimates of winds and humidity over data-rich regions (from conventional measurements) are used to show the robustness of the data and its value over regions which are currently poorly sampled.

Jedlovce, Gary J.↗

Enhancements of Bayesian Blocks; Application to Large Light Curve Databases

Bayesian Blocks are optimal piecewise linear representations (step function fits) of light-curves. The simple algorithm implementing this idea, using dynamic programming, has been extended to include more data modes and fitness metrics, multivariate analysis, and data on the circle (Studies in Astronomical Time Series Analysis. VI. Bayesian Block Representations, Scargle, Norris, Jackson and Chiang 2013, ApJ, 764, 167), as well as new results on background subtraction and refinement of the procedure for precise timing of transient events in sparse data. Example demonstrations will include exploratory analysis of the Kepler light curve archive in a search for "star-tickling" signals from extraterrestrial civilizations. (The Cepheid Galactic Internet, Learned, Kudritzki, Pakvasa1, and Zee, 2008, arXiv: 0809.0339; Walkowicz et al., in progress).

Kepler light curve archive↗

Storm Chasing from Space: Detecting severe weather phenomena from satellite platforms

Severe weather is an awe-inspiring phenomenon that affects the entire globe. Lightning, hail, damaging wind and tornadoes pose threats to society and challenges to the scientific community. Severe weather is annually responsible for tens of billions of dollars in insured losses to property, infrastructure, and agriculture. Satellite platforms offer a globally uniform approach to observing weather phenomena in remote or data-sparse regions and over the oceans. Severe convection exhibits distinct signatures in spaceborne datasets that we use to analyze severe storms. We leverage these signatures create climatologies, improve prediction, and provide a method of detection around the globe where traditional ground-based data (such as ground-based radar or human-spotter reports) are inconsistent or unavailable. Satellites in low-earth, sun-synchronous, and geostationary orbit provide a consistent, global view of severe weather from which we can examine the current global distribution, frequency, and severity of severe storms and establish a baseline to assess their future trend in a changing Earth system.

Sarah D. Bang↗