Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data driven modeling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

FluxSat: Long-term Earth Science Data Record (ESDR) for Terrestrial Gross Primary Production (GPP) based on satellite data calibrated with eddy covariance data

Gross primary production (GPP), the amount of carbon dioxide (CO 2 ) assimilated by plants through photosynthesis, is one of the most variable and uncertain components of the global carbon cycle. Global GPP has been estimated with a number of process-based models, data-driven, and hybrid approaches. Dynamic global vegetation models (DGVMs), driven by observed environmental changes, are used for global carbon budget assessments and long-term (climate) prediction. Benchmarking these and other models globally with data-driven GPP estimates is critical for understanding the land sink and ensuring accurate forecasts of the carbon cycle. In addition, global data-driven GPP estimates are crucial for studies of interannual variability, including trends that are linked to mechanisms with large uncertainties, such as the indirect CO 2 fertilization effect related to greening. In response to a community need for a GPP data set that well captures spatio-temporal variability, we developed FluxSat, a data-driven approach that optimizes the use of satellite reflectance data from the NASA MODerate-resolution Imaging Spectroradiometer (MODIS) on the Terra and Aqua satellites, calibrated using ground-based eddy covariance (EC) data. We are enhancing (spatially, higher resolution) and extending FluxSat (in time, with additional sensors) to create a high quality long term GPP Earth System Data Record (ESDR) for use in model benchmarking, carbon cycle modeling, and studies of trends and interannual variability. Our team’s objectives are to: 1. Update and document the current MODIS FluxSat GPP (daily, 0.05o and 0.5o resolutions) products with latest available MODIS and EC data sets; 2. Extend FluxSat GPP record forward in time with the Visible Infrared Imaging Radiometer Suite (VIIRS) on operational weather satellites going forward; 3. Extend FluxSat GPP record backward in time using the Advanced Very High Resolution Radiometer (AVHRR) on weather satellites dating back to 1981; 4. Provide higher spatial resolution MODIS and VIIRS GPP (0.0083o). 5. Thoroughly evaluate all FluxSat products with independent data; and 6. Create a homogenized long-term GPP record spanning 40+ years. We will discuss plans for this long-term data set that is supported through the NASA Making Earth System Data Records for Use in Research Environments (MEaSUREs) program.

gross Primary Production↗

The Distribution and Abundance of Bird Species: Towards a Satellite, Data Driven Avian Energetics and Species Richness Model

This paper addresses the fundamental question of why birds occur where and when they do, i.e., what are the causative factors that determine the spatio-temporal distributions, abundance, or richness of bird species? In this paper we outline the first steps toward building a satellite, data-driven model of avian energetics and species richness based on individual bird physiology, morphology, and interaction with the spatio-temporal habitat. To evaluate our model, we will use the North American Breeding Bird Survey and Christmas Bird Count data for species richness, wintering and breeding range. Long term and current satellite data series include AVHRR, Landsat, and MODIS.

Smith, James A.↗

Prediction of Pushback Times and Ramp Taxi Times for Departures at Charlotte Airport

When optimizing the takeoff sequence and schedule for departures at busy airports, it is important to accurately predict the taxi times from gate to runway because those are used to calculate the earliest possible takeoff times. Several airports like Charlotte Douglas International Airport show relatively long taxi times inside the ramp area with large variations, with respect to the travel times in the airport movement area. Also, the pushback process times have not been accurately modeled so far mainly due to the lack of accurate data. The recent deployment of the integrated arrival, departure, and surface traffic management system at Charlotte airport by NASA enables more accurate flight data in the airport surface operations to be obtained. Taking advantage of this system, actual pushback times and ramp taxi times from historical flight data at this airport are analyzed. Based on the analysis, a simple, data-driven prediction model is introduced for estimating pushback times and ramp transit times of individual departure flights. To evaluate the performance of this prediction model, several machine learning techniques are also applied to the same dataset. The prediction results show that the data-driven prediction model is as good as the machine learning algorithms when comparing various prediction performance metrics.

Lee, Hanbong↗

Prediction of Pushback Times and Ramp Taxi Times for Departures at Charlotte Airport

When optimizing the takeoff sequence and schedule for departures at busy airports, it is important to accurately predict the taxi times from gate to runway because those are used to calculate the earliest possible takeoff times. Several airports like Charlotte Douglas International Airport show relatively long taxi times inside the ramp area with large variations, with respect to the travel times in the airport movement area. Also, the pushback process times have not been accurately modeled so far mainly due to the lack of accurate data. The recent deployment of the integrated arrival, departure, and surface traffic management system at Charlotte airport by NASA enables more accurate flight data in the airport surface operations to be obtained. Taking advantage of this system, actual pushback times and ramp taxi times from historical flight data at this airport are analyzed. Based on the analysis, a simple, data-driven prediction model is introduced for estimating pushback times and ramp transit times of individual departure flights. To evaluate the performance of this prediction model, several machine learning techniques are also applied to the same dataset. The prediction results show that the data-driven prediction model is as good as the machine learning algorithms when comparing various prediction performance metrics.

airport surface operations↗

Testing the Accuracy of Data-Driven MHD Simulations of Active Region Evolution

Models for the evolution of the solar coronal magnetic field are vital for understanding solar activity, yet the best measurements of the magnetic field lie at the photosphere, necessitating the development of coronal models which are "data-driven" at the photosphere. We present an investigation to determine the feasibility and accuracy of such methods. Our validation framework uses a simulation of active region (AR) formation, modeling the emergence of magnetic flux from the convection zone to the corona, as a ground-truth data set, to supply both the photospheric information and to perform the validation of the data-driven method. We focus our investigation on how the accuracy of the data-driven model depends on the temporal frequency of the driving data. The Helioseismic and Magnetic Imager on NASA's Solar Dynamics Observatory produces full-disk vector magnetic field measurements at a 12-minute cadence. Using our framework we show that ARs that emerge over 25 hours can be modeled by the data-driving method with only approximately 1 percent error in the free magnetic energy, assuming the photospheric information is specified every 12 minutes. However, for rapidly evolving features, under-sampling of the dynamics at this cadence leads to a strobe effect, generating large electric currents and incorrect coronal morphology and energies. We derive a sampling condition for the driving cadence based on the evolution of these small-scale features, and show that higher-cadence driving can lead to acceptable errors. Future work will investigate the source of errors associated with deriving plasma variables from the photospheric magnetograms as well as other sources of errors, such as reduced resolution, instrument bias, and noise.

Leake, James E.↗

Remote Sensing-Driven Hydrodynamic Modeling in Data-Scarce Regions: Integrating ICESat-2, Sentinel-2, SWOT and Re-analysis Models for Coastal Monitoring

Hydrodynamic models in coastal and estuarine systems are typically constrained by sparse bathymetry, boundary, and validation data, especially in regions where field campaigns are costly or impractical. Here we develop and test a fully satellite-driven framework for hydrodynamic modeling in South Africa’s Langebaan Lagoon without using any local in situ measurements. Bathymetry is derived by training multispectral Sentinel-2 reflectance against ICESat-2 ATL24 photon-derived depths using an XGBoost model optimized with Bayesian search. The final satellite-derived bathymetry reproduces independent ATL24 points with RMSE = 0.45 m and R 2 = 0.97. This bathymetry was used in a depth-averaged Delft3D Flexible Mesh model driven at the open boundary by TPXO tidal harmonics and by ERA5 winds. We validate modeled water surface elevation against 16 SWOT low-rate (250 m, unsmoothed) passes in 2023. SWOT–model comparisons yield an overall RMSE of 0.11 m and R 2 = 0.61, with typical point differences <0.10 m (∼7% of the 1.5 m tidal range), and showed consistent spatial gradients in water level from the offshore boundary, through Saldanha Bay, and into the lagoon. At the offshore boundary, TPXO and SWOT sea surface heights agree closely (R 2 = 0.86). A simple phase adjustment of ∼26,min between TPXO and SWOT lowers the RMSE from 0.18,m to 0.11,m, showing that phase offset accounts for some of the discrepancy, with additional errors likely linked to non-tidal signals. Our results demonstrate that combining passive optical, photon-counting LiDAR, radar interferometry, and global tidal/atmospheric models enables robust, transferrable hydrodynamic modeling in data-scarce coastal systems, offering a cost-effective pathway for monitoring.

ICESat-2↗

Investigation and Development of Data-Driven D-Region Model for HF Systems Impacts

Space Environment Corporation (SEC) and RP Consultants (RPC) are to develop and validate a weather-capable D region model for making High Frequency (HF) absorption predictions in support of the HF communications and radar communities. The weather-capable model will assimilate solar and earth space observations from NASA satellites. The model will account for solar-induced impacts on HF absorption, including X-rays, Solar Proton Events (SPE's), and auroral precipitation. The work plan includes: I . Optimize D-region model to quickly obtain ion and electron densities for proper HF absorption calculations. 2. Develop indices-driven modules for D-region ionization sources for low, mid, & high latitudes including X-rays, cosmic rays, auroral precipitation, & solar protons. (Note: solar spectrum & auroral modules already exist). 3. Setup low-cost monitors of existing HF beacons and add one single-frequency beacon. 4. Use PENEX HF-link database with HF monitor data to validate D-region/HF absorption model using climatological ionization drivers. 5. Develop algorithms to assimilate NASA satellite data of solar, interplanetary, and auroral observations into ionization source modules. 6. Use PENEX HF-link & HF-beacon data for skill score comparison of assimilation versus climatological D-region/HF absorption model. Only some satellites are available for the PENEX time period, thus, HF-beacon data is necessary. 7. Use HF beacon monitors to develop HF-link data assimilation algorithms for regional improvement to the D-region/HF absorption model.

Eccles, J. V.↗

Strategies for concurrent processing of complex algorithms in data driven architectures

Performance modeling and performance enhancement for periodic execution of large-grain, decision-free algorithms in data flow architectures are discussed. Applications include real-time implementation of control and signal processing algorithms where performance is required to be highly predictable. The mapping of algorithms onto the specified class of data flow architectures is realized by a marked graph model called algorithm to architecture mapping model (ATAMM). Performance measures and bounds are established. Algorithm transformation techniques are identified for performance enhancement and reduction of resource (computing element) requirements. A systematic design procedure is described for generating operating conditions for predictable performance both with and without resource constraints. An ATAMM simulator is used to test and validate the performance prediction by the design procedure. Experiments on a three resource testbed provide verification of the ATAMM model and the design procedure.

Som, Sukhamoy↗

Fault Diagnosis of Power Components with Reliability Assessment in Extraterrestrial Microgrids

This research investigates the possible failures caused by aging and other environmental and external factors that could significantly impact the performance of extraterrestrial power systems. Additionally, it presents a reliability assessment model for the space microgrid based on fault tree analysis (FTA). The reliability assessment model developed in this paper represents a tool that can be used by engineers to harden the system design for operational and economic benefits. To improve the reliability of the system, this work provides a broad review of the different fault detection and diagnosis (FDD) algorithms used for power microgrids and space applications. Using data sets from the Habitat Simulator developed through the NASA-funded Resilient Extraterrestrial Habitat Institute, this paper compares the applicability and accuracy of the different FDD methods. The primary FDD approach proposed and assessed in this work is based on the Markov reliability model. It predicts and detects future faults in the space microgrids by using past data samples and categorizing them into different classes. Data-driven-based models such as artificial neural networks are also investigated, tested, and evaluated using simulation data sets. According to the simulation results and the broad FDD algorithm comparison, this study provides the crew or maintenance engineers with a clear methodology to detect and localize power system failures.

Leila Chebbo↗

A Data-driven, Physics-based Transport Model of Solar Energetic Particles Accelerated by Coronal Mass Ejection Shocks Propagating through the Solar Coronal and Heliospheric Magnetic Fields

In an effort to develop computational tools for predicting radiation hazards from solar energetic particles (SEPs), we have created a data-driven physics-based particle transport model to calculate the injection, acceleration, and propagation of SEPs from coronal mass ejection (CME) shocks traversing through the solar corona and interplanetary magnetic fields. The model runs on an input of corona and heliospheric plasma and magnetic field configuration from a magnetohydrodynamic model driven by solar photospheric magnetic field measurements superposed with observed CME shocks determined from coronagraph images. SEP source particles are injected at the shock using the result of diffusive shock acceleration formulation from a characteristic obliquity-dependent injection from a heated solar wind thermal tail population. With several advanced computation techniques involving stochastic simulation and integration, the model obtains the particle intensity at any location in interplanetary space through the rigorous solution to the time-dependent 5D focus transport equation in the phase space that includes perpendicular diffusion. We apply the model to the 2011 November 3 CME event. The calculation results reproduce multispacecraft SEP observations at Earth and STEREO-B reasonably well without normalization of particle flux. The observations at STEREO-A can be reproduced by rescaling particle energy or modified energy dependence of particle diffusion coefficients. This circumsolar SEP event seen by spacecraft at Earth, STEREO-A, and STEREO-B at widely separated longitudes can be explained by diffusive shock acceleration by a single CME shock with a moderate speed.

Solar energetic particles↗

Application of Machine Learning Techniques to Aviation Operations: NASA Case Studies

There is an increasing interest in applying methods based on Machine Learning Techniques(MLT) to problems in aviation operations. The current interest is based on developments in Cloud Computing, the availability of open software and the success of MLT in automation, consumer behavior and finance involving large database. Historically aviation operations have been analyzed using physics-based models and provide information for making operational decisions. This talk describes issues to be addressed in applying either model-driven or data-driven methods. Aviation operations involving many decision makers, multiple objectives, poor or unavailable physics-based models and a rich historical database are prime candidates for analysis using data-driven methods. The issues relating to data, feature selection and validation of the models are illustrated by examining case studies of the application of MLT to problems in air traffic management at NASA. Further research is needed in the application of MLT to critical aviation operations. As always, the best approach depends on the task, the physical understanding of the problem and the quality and quantity of the available data.

Sridhar, Banavar↗

Development of Solar Wind Model Driven by Empirical Heat Flux and Pressure Terms

We are developing a time stationary self-consistent 2D MHD model of the solar corona and solar wind as suggested by Sittler et al. (2003). Sittler & Guhathakurta (1999) developed a semiempirical steady state model (SG model) of the solar wind in a multipole 3-streamer structure, with the model constrained by Skylab observations. Guhathakurta et al. (2006) presented a more recent version of their initial work. Sittler et al. (2003) modified the SG model by investigating time dependent MHD, ad hoc heating term with heat conduction and empirical heating solutions. Next step of development of 2D MHD models was performed by Sittler & Ofman (2006). They derived effective temperature and effective heat flux from the data-driven SG model and fit smooth analytical functions to be used in MHD calculations. Improvements of the Sittler & Ofman (2006) results now show a convergence of the 3-streamer topology into a single equatorial streamer at altitudes > 2 R(sub S). This is a new result and shows we are now able to reproduce observations of an equatorially confined streamer belt. In order to allow our solutions to be applied to more general applications, we extend that model by using magnetogram data and PFSS model as a boundary condition. Initial results were presented by Selwa et al. (2008). We choose solar minimum magnetogram data since during solar maximum the boundary conditions are more complex and the coronal magnetic field may not be described correctly by PFSS model. As the first step we studied the simplest 2D MHD case with variable heat conduction, and with empirical heat input combined with empirical momentum addition for the fast solar wind. We use realistic magnetic field data based on NSO/GONG data, and plan to extend the study to 3D. This study represents the first attempt of fully self-consistent realistic model based on real data and including semi-empirical heat flux and semi-empirical effective pressure terms.

Sittler, Edward C., Jr.↗

A Robust Machine Learning Schema for Developing, Maintaining, and Disseminating Machine Learning Models

Recent advances in the development of machine learning (ML) algorithms have enabled the creation of predictive models that can improve decision making, decrease computational cost, and improve efficiency in a variety of fields. As an organization begins to develop and implement such models, the data used in the training, validation, and testing of ML models, the model parameters, and the use cases or limitations of the models must be properly stored to ensure models are both fully traceable and used correctly. In the context of predicting material behavior, advances in computationally intense, physics-based modeling of material behavior at various length scales and the emergence of Integrated Computational Materials Engineering (ICME) have driven the need for developing data-driven surrogate models of the physics-based simulation tools using ML techniques. Surrogate model development allows for accurate material behavior prediction at a fraction of the cost of its physics-based counterpart, allowing for multiscale simulations of real-world applications, further enabling the ability to design fit-for-purpose materials for a reasonable computational investment. However, training such models requires extensive data, and thus, effective data management is necessary to reach the full potential that ML can offer to material design and ICME. This paper proposes a generalized, robust schema that allows organizations to store both real (experimental) and virtual (simulation) data used to train ML models and the defining model parameters and architectures within the Granta MI Platform. The developed schema allows for various types of data inputs and outputs, including single point values, time-series data, and images that can be used in the prediction of material behavior, while following outlined best practices for effective data management. An effective schema for ML data and models can help prevent the recreation of virtual/real training data and surrogate models, help reduce the time to create new models similar to existing ones by offering a starting point in the hyperparameter determination stages, minimize resources devoted to verification and validation (V&V) and certification of models, and ensure that data and surrogate models are not misused due to full traceability of both the data and ML model. It also allows organizations access to models that have already been developed, such that they can be used in the design of new materials, enabling the overall goals of ICME.

Brandon L. Hearley↗

Predicting Fiber Failure of Plain Weave Fabric with Recursive Multiscale Micromechanics

Recent advances in the development of machine learning (ML) algorithms have enabled the creation of predictive models that can improve decision making, decrease computational cost, and improve efficiency in a variety of fields. As an organization begins to develop and implement such models, the data used in the training, validation, and testing of machine learning models, the model parameters, and the use cases or limitations of the models must be properly stored to ensure models are both fully traceable and used correctly. In the context of predicting material behavior, advances in computationally intense, physics-based, modeling of material behavior at various length scales, and the emergence of Integrated Computational Materials Engineering (ICME) have driven the need for developing data-driven surrogate models of the physics-based simulation tools using machine learning (ML) techniques. Surrogate model development allows for accurate material behavior prediction at a fraction of the cost of its physics-based counterpart, allowing for multiscale simulations of real-world applications, further enabling the ability to design fit-for-purpose materials for a reasonable computational investment. However, training such models requires extensive data, and thus effective data management is necessary to reach the full potential that ML can offer to material design and ICME. This paper proposes a generalized, robust schema that allows organizations to store both real (experimental) and virtual (simulation) data used to train machine learning models and the defining model parameters and architectures. The developed schema allows for various types of data inputs and outputs, including single point values, time-series data, and images that can be used in for various types of machine learning models while following outlined best practices for effective data management. An effective schema for machine learning data and models can help prevent the recreation of virtual/real training data and surrogate models, can help reduce the time to create new models similar to existing ones by offering a starting point in the hyperparameter determination stages, minimize resources devoted to verification and validation (V&V) and certification of models, and ensure that data and surrogate models are not misused due to full traceability of both the data and ML model. It also allows organizations access to models that have already been developed, such that they can be used in the design of new materials, enabling the overall goals of ICME.

Failure↗