Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Sparse Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Covariate Dependent Sparse Functional Data Analysis

This study proposes a method to incorporate covariate information into sparse functional data analysis. The method aims at cases where each subject has a limited number of longitudinal measurements and is associated with static covariates. This research is motivated by several use cases in practice. One representative example is void swelling, a nuclear-specific material degradation mechanism. Void swelling is affected by many covariates, including alloy composition and irradiation type. How to accurately model the complicated joint effects of such covariates on the swelling process is the key to mitigating the effect of swelling and ensuring safe operation. Unlike most of the existing methods, the proposed method can handle high-dimensional covariates with the informative covariate identification procedure and sparse and irregularly spaced measurements, that is, does not require complete or dense observations. The main innovation of the proposed method is that we model the variation coming from covariates and the variation left conditioned on covariates, such that the functional principal component analysis and Gaussian process can be conducted in a unified manner. Further, we also propose a systematic approach to identify important covariates in the hypothesis testing context. The methodology is demonstrated on applications in nuclear engineering and healthcare and simulation studies.

42 ENGINEERING↗

Image Reconstruction from Sparse-view Data Acquired with Portable X-ray Devices

• Portable X-ray systems enable on-site 3D imaging for non-invasive inspection of suspicious packages and explosives. • Existing reconstruction algorithms (e.g., FDK or Feldkamp, Davis and Kress) require hundreds of projections over 360 degrees. • Sparse-view scan reduces scanning time and setup effort, making it ideal for field use in timecritical scenarios. • Existing reconstruction algorithms introduce severe artifacts when applied to sparse-view data. • We developed a total variation (TV)-based optimization algorithm for yielding 3D images from sparse-view data collected with our portable X-ray imaging system.

Xia, Dan [University of Chicago, Chicago, IL]↗

High-Fidelity Heavy-Duty Vehicle Modeling Using Sparse Telematics Data

Heavy-duty commercial vehicles consume a significant amount of energy due to their large size and mass, directly leading to vehicle operators prioritizing energy efficiency to reduce operational costs and comply with environmental regulations. One tool that can be used for the evaluation of energy efficiency in heavy-duty vehicles is the evaluation of energy efficiency using vehicle modeling and simulation. Simulation provides a path for energy efficiency improvement by allowing rapid experimentation of different vehicle characteristics on fuel consumption without the need for costly physical prototyping. The research presented in this paper focuses on using real-world, sparsely sampled telematics data from a large fleet of heavy-duty vehicles to create high-fidelity models for simulation. Samples in the telematics dataset are collected sporadically, resulting in sparse data with an infrequent and irregular sampling rate. Captured in the dataset was geospatial information, time series measurements, and vehicle-specific metadata from a subset of 96 vehicles from varied geographic regions across North America. A series of custom algorithms was developed to process vehicle data and derive both vehicle model input parameters and representative drive cycles. Derived models provide a basis on which to simulate real-world vehicles and iterate on vehicle aerodynamics, auxiliary power loads, transmission shift schedules, and other parameters to achieve reduced fuel consumption and increase energy efficiency. Notably, these models were developed without the use of expensive field data collection, using only data collected through fleet telematics. Processed representative drive cycles are used to validate the fuel economy of derived models. The models developed through this research allow for more representative vehicle simulations with increased flexibility regarding vehicle-to-vehicle variations.

ADVANCED PROPULSION SYSTEMS↗

Restoring the discontinuous heat equation source using sparse boundary data and dynamic sensors

Abstract This study focuses on addressing the inverse source problem associated with the parabolic equation. We rely on sparse boundary flux data as our measurements, which are acquired from a restricted section of the boundary. While it has been established that utilizing sparse boundary flux data can enable source recovery, the presence of a limited number of observation sensors poses a challenge for accurately tracing the inverse quantity of interest. To overcome this limitation, we introduce a sampling algorithm grounded in Langevin dynamics that incorporates dynamic sensors to capture the flux information. Furthermore, we propose and discuss two distinct dynamic sensor migration strategies. Remarkably, our findings demonstrate that even with only two observation sensors at our disposal, it remains feasible to successfully reconstruct the high-dimensional unknown parameters.

Mathematics↗

Transfer-Learnt Energy Models for Predicting Electricity Consumption in Buildings with Limited and Sparse Field Data

Modeling energy consumption is critical for energy-efficient utilization of the electric appliances in a building, smart grid programs (like demand-response), and many other smart home applications. State-of-the-art energy modeling techniques either rely on theoretical models, or extensive instrumentation of the building envelope to gather ``big" data to train a deep neural network. While theoretical models are often limited by their estimation accuracy, it is not always feasible to gather a significant amount of field data. In this paper, we explore transfer learning-based strategies to train much more accurate model for energy estimation when using a sparse field data. We transferred knowledge, in the form of data and parameters, from the simulation framework to the field data. We evaluated the efficacy of our approach on field data collected from six commercial buildings and our results indicate that transfer learning-based models trained over one month data can perform comparative (and in some cases better) than the state-of-the-art machine learning and deep learning solutions.

Jain, Milan↗

Towards Characterizing the Variability in the Loading Demands of an Unmanned Aerial Vehicle

This paper presents a computational methodology to characterize and quantify the variability in the power demands during the take-off of an unmanned aerial vehicle (UAV). A lithium-ion battery-based power system is used to power the unmanned aerial vehicle, and the capabilities of the unmanned aerial vehicle are driven by the amount of charge in this battery. In order to design the power system, it is necessary to analyze the power and charge requirements of the UAV. This paper focuses on the take-off segment, and aims to quantify the amount of charge that is required for this particular segment. Sparse data is available through different flight tests and this data is used to analyze the flight profile and the charge requirement during take-off. The amount of charge required for take-off depends on several factors that are not only variable but cannot be controlled in reality, and hence, the entire flight profile and the corresponding charge requirement are variable in nature. The information available through flight tests is converted into multi-dimensional sparse data and a new method is developed in this paper for variability characterization using multi-dimensional sparse data. This analysis is useful for prognostics and health management where it is necessary to anticipate future charge requirements in order to compute the end-of-discharge of the battery, and hence, the remaining useful life of the power system.

unmanned aerial vehicle↗

A Computational Information Criterion for Particle-Tracking with Sparse or Noisy Data

Traditional probabilistic methods for the simulation of advection-diffusion equations (ADEs) often overlook the entropic contribution of the discretization, e.g., the number of particles, within associated numerical methods. Many times, the gain in accuracy of a highly discretized numerical model is outweighed by its associated computational costs or the noise within the data. Herein, we address the question of how many particles are needed in a simulation to best approximate and estimate parameters in one-dimensional advective-diffusive transport. To do so, we use the well-known Akaike Information Criterion (AIC) and a recently-developed correction called the Computational Information Criterion (COMIC) to guide the model selection process. Random-walk and mass-transfer particle tracking methods are employed to solve the model equations at various levels of discretization. Numerical results demonstrate that the COMIC provides an optimal number of particles that can describe a more efficient model in terms of parameter estimation and model prediction compared to the model selected by the AIC even when the data is sparse or noisy, the sampling volume is not uniform throughout the physical domain, or the error distribution of the data is non-IID Gaussian.

97 MATHEMATICS AND COMPUTING↗

High-performance equation solvers and their impact on finite element analysis

The role of equation solvers in modern structural analysis software is described. Direct and iterative equation solvers which exploit vectorization on modern high-performance computer systems are described and compared. The direct solvers are two Cholesky factorization methods. The first method utilizes a novel variable-band data storage format to achieve very high computation rates and the second method uses a sparse data storage format designed to reduce the number of operations. The iterative solvers are preconditioned conjugate gradient methods. Two different preconditioners are included; the first uses a diagonal matrix storage scheme to achieve high computation rates and the second requires a sparse data storage scheme and converges to the solution in fewer iterations that the first. The impact of using all of the equation solvers in a common structural analysis software system is demonstrated by solving several representative structural analysis problems.

Poole, Eugene L.↗

High-performance equation solvers and their impact on finite element analysis

The role of equation solvers in modern structural analysis software is described. Direct and iterative equation solvers which exploit vectorization on modern high-performance computer systems are described and compared. The direct solvers are two Cholesky factorization methods. The first method utilizes a novel variable-band data storage format to achieve very high computation rates and the second method uses a sparse data storage format designed to reduce the number od operations. The iterative solvers are preconditioned conjugate gradient methods. Two different preconditioners are included; the first uses a diagonal matrix storage scheme to achieve high computation rates and the second requires a sparse data storage scheme and converges to the solution in fewer iterations that the first. The impact of using all of the equation solvers in a common structural analysis software system is demonstrated by solving several representative structural analysis problems.

Poole, Eugene L.↗

Emulation Modeling for Development of Cyber-Defense Capabilities for Satellite Systems

The objective of this project was to develop a novel capability to generate synthetic data sets for the purpose of training Machine Learning (ML) algorithms for the detection of malicious activities on satellite systems. The approach experimented with was to a) generate sparse data sets using emulation modeling and b) enlarge the sparse data using Generative Adversarial Networks (GANs). We based our emulation modeling on the Open Source NASA Operational Simulator for Small Satellites (NOS3) developed by the Katherine Johnson Independent Verification and Validation (IV&V) program in West Virginia. Significant new capabilities on NOS3 had to be developed for our data set generation needs. To expand these data sets for the purpose of training ML, we experimented with a) Extreme Learning Machines (ELMs) and b) Wasserstein-GANs (WGAN-GP).

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

A DATA EFFICIENT SPARSE MODELING FRAMEWORK FOR POWER ESTIMATION IN WATER TREATMENT SENSING OPERATIONS

With increasing freshwater scarcity, advanced process design mechanisms such as Closed-Circuit Reverse Osmosis (CCRO) and Digital/Physical Twin systems are gaining traction in water treatment and reuse operations. While digital and physical twin models enable improved system insight and control, their development is often expensive and computationally intensive, requiring large volumes of synthetic or experimental data to characterize underlying process dynamics. This work introduces a sparse surrogate modeling framework to estimate power consumption from measured flow and pressure variables, along with their nonlinear polynomial and interaction expansions. To ensure model reliability and reduce overfitting, a two-stage pipeline is proposed. First, a dynamic data filtering algorithm is employed to remove uninformative observations and transient operational states. Second, a sparse penalized regression technique is applied to select a minimal set of parsimonious features. The proposed model achieves high sparsity, retaining only 7 out of 34 candidate features (≈79.41% sparsity) while delivering a root mean square error (RMSE) of 0.072 on the test dataset.

Mukherjee, Subrata [ORNL] (ORCID:0000000309930338)↗

A Comprehensive Validation Methodology for Sparse Experimental Data

A comprehensive program of verification and validation has been undertaken to assess the applicability of models to space radiation shielding applications and to track progress as models are developed over time. The models are placed under configuration control, and automated validation tests are used so that comparisons can readily be made as models are improved. Though direct comparisons between theoretical results and experimental data are desired for validation purposes, such comparisons are not always possible due to lack of data. In this work, two uncertainty metrics are introduced that are suitable for validating theoretical models against sparse experimental databases. The nuclear physics models, NUCFRG2 and QMSFRG, are compared to an experimental database consisting of over 3600 experimental cross sections to demonstrate the applicability of the metrics. A cumulative uncertainty metric is applied to the question of overall model accuracy, while a metric based on the median uncertainty is used to analyze the models from the perspective of model development by analyzing subsets of the model parameter space.

Norman, Ryan B.↗

A Global Assessment of Added Value in the SMAP Level-4 Soil Moisture Product Relative to Its Baseline Land Surface Model

The Soil Moisture Active Passive (SMAP) Level-4 product provides enhanced soil moisture estimates by assimilating SMAP brightness temperature observations into a land surface model. Here, an unbiased qualitative estimate of the relative skill of SMAP Level-4 and model-only surface soil moisture (versus true soil moisture) is derived using only one additional noisy (but independent) soil moisture product. The method is applied globally and verified using high-quality, ground-based measurements where available. Results demonstrate that assimilating SMAP brightness temperature has relatively little impact in data-rich areas like the United States and Europe. In contrast, much larger improvement is observed in data-sparse regions, including much of Africa and central Australia, where model-only simulations are disproportionately impacted by low-quality model forcing. Therefore, ground validation conducted in data-rich areas does not adequately sample the added value of SMAP data assimilation for data-sparse regions and substantially underestimates the added skill provided by the SMAP Level-4 system.

SMAP L4↗

Design of Materials with Alchemite

Machine learning models that establish the relationships between materials processing and properties can enable inverse design of materials through active learning. Alchemite is a commercial software that can perform inverse materials design on sparse data. Here we evaluate Alchemite’s performance on a dataset of shape memory alloys and a dataset of heat exchangers compared to baseline random forest models. Alchemite had higher accuracy when making predictions on sparse data and was more accurate or nearly as accurate as random forests on complete datasets while also quantifying uncertainty. The software was also used to suggest processing steps and design parameters to optimize properties and performance; however, physical validation of the suggested design parameters was beyond the scope of this work. Several useful design insights were gained about the impact of the design parameters on properties and performance including the importance of dopant choice and amount for shape memory alloys and the importance of height and weight on the thermal resistance of heat exchangers.

Machine learning↗

Data for A Hybrid Biophysical-Machine Learning Framework for Diurnal Surface Energy Flux Estimation Using Proximal Sensing

Thermal infrared-based remote sensing of surface energy fluxes has traditionally relied on high spatial resolution satellite data with revisit frequencies on the order of weeks. In this study, we evaluate a biophysics-based analytical surface energy balance model for predicting latent energy (LE) and sensible heat (H) fluxes using proximal sensing observations. The Surface Temperature Initiated Closure (STIC1.2) model has been extensively validated across a wide range of spatial and temporal scales using various satellite-derived thermal infrared data sets. Here we extend this validation by applying STIC at sub-hourly temporal resolution over multiple growing seasons for four distinct agricultural systems. We further develop and evaluate novel STIC variants that incorporate machine learning (ML) techniques to eliminate the need for surface energy balance observations, specifically net radiation and soil heat flux, thereby enhancing model applicability in data-sparse settings. The integration of a ML component to estimate surface available energy is shown to have strong predictive performance for both LE (R2 = 0.81–0.94) and H (R2 = 0.46–0.72) across all agricultural systems examined here, demonstrating the potential of hybrid biophysical-machine learning approaches for surface energy balance modeling with minimal data requirements. This study concludes with a novel application of explainable machine learning (exML) to diagnose sources of model error. This exML framework attributes residual prediction errors to both model input variables and environmental drivers not explicitly included in the simulation experiments. This approach provides a new pathway for improving model design and integrating previously overlooked yet influential variables into future model iterations.

AI/ML↗

Three-dimensional reconstruction of x-ray emission volumes in magnetized liner inertial fusion from sparse projection data using a learned basis

The ability to visualize x-ray and neutron emission from fusion plasmas in 3D is critical to understand the origin of the complex shapes of the plasmas in experiments. Unfortunately, this remains challenging in experiments that study a fusion concept known as Magnetized Liner Inertial Fusion (MagLIF) due to a small number of available diagnostic views. Here, we present a basis function-expansion approach to reconstruct MagLIF stagnation plasmas from a sparse set of x-ray emission images. A set of natural basis functions is “learned” from training volumes containing quasi-helical structures whose projections are qualitatively similar to those observed in experimental images. Tests on several known volumes demonstrate that the learned basis outperforms both a cylindrical harmonic basis and a simple voxel basis with additional regularization, according to several metrics. Two-view reconstructions with the learned basis can estimate emission volumes to within 11% and those with three views recover morphology to a high degree of accuracy. The technique is applied to experimental data, producing the first 3D reconstruction of a MagLIF stagnation column from multiple views, providing additional indications of liner instabilities imprinting onto the emitting plasma.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Processing Aleatory and Epistemic Uncertainties in Experimental Data From Sparse Replicate Tests of Stochastic Systems for Real-Space Model Validation

This paper presents a practical methodology for propagating and processing uncertainties associated with random measurement and estimation errors (that vary from test-to-test) and systematic measurement and estimation errors (uncertain but similar from test-to-test) in inputs and outputs of replicate tests to characterize response variability of stochastically varying test units. Also treated are test condition control variability from test-to-test and sampling uncertainty due to limited numbers of replicate tests. These aleatory variabilities and epistemic uncertainties result in uncertainty on computed statistics of output response quantities. The methodology was developed in the context of processing experimental data for “real-space” (RS) model validation comparisons against model-predicted statistics and uncertainty thereof. The methodology is flexible and sufficient for many types of experimental and data uncertainty, offering the most extensive data uncertainty quantification (UQ) treatment of any model validation method the authors are aware of. It handles both interval and probabilistic uncertainty descriptions and can be performed with relatively little computational cost through use of simple and effective dimension- and order-adaptive polynomial response surfaces in a Monte Carlo (MC) uncertainty propagation approach. A key feature of the progressively upgraded response surfaces is that they enable estimation of propagation error contributed by the surrogate model. Sensitivity analysis of the relative contributions of the various uncertainty sources to the total uncertainty of statistical estimates is also presented. Finally, the methodologies are demonstrated on real experimental validation data involving all the mentioned sources and types of error and uncertainty in five replicate tests of pressure vessels heated and pressurized to failure. Simple spreadsheet procedures are used for all processing operations.

97 MATHEMATICS AND COMPUTING↗

Development of a Framework and Methodology for an Advanced Reactor Materials Environmental Effects Design Guide

Advanced non-light-water reactor components may operate at elevated temperature while experiencing cyclic loading, significant neutron irradiation, and exposure to reactor coolant. ASME Boiler and Pressure Vessel Code, Section III, Division 5, provides design rules for elevated-temperature service but does not include specific procedures to account for environmental effects on material properties. This report develops an initial framework and methodology for an Environmental Effects Design Guide (EEDG) focused on neutron irradiation; coolant-environment effects are reserved for future work. The proposed approach treats irradiation as a property-based overlay on the existing Division 5 design process, with two routes: a sparse-data route applying two reduction factors — FCR on creep-rupture strength and FF on fatigue life — for the creep-fatigue evaluations that typically control the design of advanced high-temperature reactor components, and a fuller framework developing the property-to-rule chain across the four Division 5 checks (primary load, strain limits and ratcheting, creep-fatigue, and buckling), together with swelling and weldments as scope items. Both routes are scoped by an in-pile qualification that restricts the use of post-irradiation-examination-derived properties in regimes where an in-pile mechanism could control the design outcome. Illustrative outputs derived on a compiled annealed Type 316 database — FCR ≈ 0.78–0.86 and FF ≈ 0.4 — demonstrate the calculation method within that specific dataset. The framework is an initial, testable design-rule concept; it identifies a practical path for preliminary design evaluations under sparse data and the material data and testing needed to develop the framework further.

Barua, Bipul (ORCID:0000000247184113)↗