Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data-driven machine learning model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

The Land Surface Data Toolkit (LDT v7.2) - A Data Fusion Environment for Land Data Assimilation Systems

The effective applications of land surface models (LSMs) and hydrologic models pose a varied set of data input and processing needs, ranging from ensuring consistency checks to more derived data processing and analytics. This article describes the development of the Land surface Data Toolkit (LDT), which is an integrated framework designed specifically for processing input data to execute LSMs and hydrological models. LDT not only serves as a preprocessor to the NASA Land Information System (LIS), which is an integrated framework designed for multi-model LSM simulations and data assimilation (DA) integrations, but also as a land-surface-based observation and DA input processor. It offers a variety of user options and inputs to processing datasets for use within LIS and stand-alone models. The LDT design facilitates the use of common data formats and conventions. LDT is also capable of processing LSM initial conditions and meteorological boundary conditions and ensuring data quality for inputs to LSMs and DA routines. The machine learning layer in LDT facilitates the use of modern data science algorithms for developing data-driven predictive models. Through the use of an object-oriented framework design, LDT provides extensible features for the continued development of support for different types of observational datasets and data analytics algorithms to aid land surface modeling and data assimilation.

droughts and floods↗

Taxi Time Prediction at Charlotte Airport Using Fast-Time Simulation and Machine Learning Techniques

Accurate taxi time prediction can be used for more efficient runway scheduling to increase runway throughput and reduce taxi times and fuel consumptions on the airport surface. This paper describes two different approaches to predicting taxi times, which are a data-driven analytical method using machine learning techniques and a fast-time simulation-based approach. These two taxi time prediction methods are applied to realistic flight data at Charlotte Douglas International Airport (CLT) and assessed with actual taxi time data from the human-in-the-loop simulation for CLT airport operations using various performance measurement metrics. Based on the preliminary results, we discuss how the taxi time prediction accuracy can be affected by the operational complexity at this airport and how we can improve the fast-time simulation model for implementing it with an airport scheduling algorithm in real-time operational environment.

Lee, Hanbong↗

Application of Machine Learning Techniques to Aviation Operations: Promises and Challenges

There is an increasing interest in applying methods based on Machine Learning Techniques (MLT) to problems in aviation operations. The current interest is based on developments in Cloud Computing, the availability of open software and the success of MLT in automation, consumer behavior and finance involving large database. This paper reviews the current-state-of-the art in applying MLT to aviation operations, its promises and challenges. Historically aviation operations have been analyzed using physics-based models and provide information for making operational decisions. This paper compares the methodology used in and issues to be addressed in applying either model-driven or data-driven methods. Some aviation operations involving many decision makers, multiple objectives, poor or unavailable physics-based models and a rich historical database are prime candidates for data-driven methods. The application of MLT to aviation operations falls into three categories: (a) based on the lack of a physics-based model, MLT is the favored approach, (b) marginal difference between regression methods using physics-based models and MLT and (c) better results using a blend of physics-based methods combined with MLT. As always, the best approach depends on the task, the physical understanding of the problem and the quality and quantity of the available data.

Sridhar, Banavar↗

Observations on the Application of Machine Learning Techniques to Aviation Operations

There is an increasing interest in applying methods based on Machine Learning Techniques (MLT) to problems in aviation operations. The current interest is based on developments in Cloud Computing, the availability of open software and the success of MLT in automation, consumer behavior and finance involving large database. Historically aviation operations have been analyzed using physics-based models and provide information for making operational decisions. This talk describes issues to be addressed in applying either model-driven or data-driven methods. Aviation operations involving many decision makers, multiple objectives, poor or unavailable physics-based models and a rich historical database are prime candidates for analysis using data-driven methods. The issues are illustrated by a detailed example and summary of current research in the area. The application of MLT to aviation operations falls into two categories: (a) based on the lack of a physics-based model, MLT is the favored approach and (b) marginal difference between regression methods using physics-based models and MLT. Further research is needed in the selection of MLT to critical aviation operations. As always, the best approach depends on the task, the physical understanding of the problem and the quality and quantity of the available data.

Sridhar, Banavar↗

Application of Machine Learning Techniques to Aviation Operations: A Case Study

There is an increasing interest in applying methods based on Machine Learning Techniques (MLT) to problems in aviation operations. The current interest is based on developments in Cloud Computing, the availability of open software and the success of MLT in automation, consumer behavior and finance involving large database. Historically aviation operations have been analyzed using physics-based models and provide information for making operational decisions. This talk describes issues to be addressed in applying either model-driven or data-driven methods. Aviation operations involving many decision makers, multiple objectives, poor or unavailable physics-based models and a rich historical database are prime candidates for analysis using data-driven methods. The issues are illustrated by a detailed example and summary of current research in the area. The application of MLT to aviation operations falls into two categories 58; (a) based on the lack of a physics-based model, MLT is the favored approach and (b) marginal difference between regression methods using physics-based models and MLT. Further research is needed in the selection of MLT to critical aviation operations. As always, the best approach depends on the task, the physical understanding of the problem and the quality and quantity of the available data.

Sridhar, Banavar↗

Application of Machine Learning Techniques to Aviation Operations: NASA Case Studies

There is an increasing interest in applying methods based on Machine Learning Techniques(MLT) to problems in aviation operations. The current interest is based on developments in Cloud Computing, the availability of open software and the success of MLT in automation, consumer behavior and finance involving large database. Historically aviation operations have been analyzed using physics-based models and provide information for making operational decisions. This talk describes issues to be addressed in applying either model-driven or data-driven methods. Aviation operations involving many decision makers, multiple objectives, poor or unavailable physics-based models and a rich historical database are prime candidates for analysis using data-driven methods. The issues relating to data, feature selection and validation of the models are illustrated by examining case studies of the application of MLT to problems in air traffic management at NASA. Further research is needed in the application of MLT to critical aviation operations. As always, the best approach depends on the task, the physical understanding of the problem and the quality and quantity of the available data.

Sridhar, Banavar↗

Multisensor Machine Learning to Retrieve High Spatiotemporal Resolution Land Surface Temperature

Climate change is making heat waves more frequent, long-lasting, and severe. While multiple satellite types provide data to monitor surface temperature, geostationary (GEO) sensors provide near-continuous, continental-scale observations which can better capture the diurnal variability of land surface temperature (LST) than intermittent observations from low-earth orbit (LEO) sensors. However, standard products from GEO satellites are available at coarsened spatial and temporal resolutions compared to the native sensor resolution. Using datasets from the NASA Earth Exchange, we leveraged co-located, co-temporal observations from LEO and GEO satellites to learn a data-driven mapping using a convolutional neural network. The resulting NASA Earth eXchange Artificial Intelligence LST (NEXAI-LST) achieved a mean absolute error of 1.73 K relative to the target LEO product and improves on both spatial and temporal resolution [2 km, 10 minute] compared to the GEO full disk standard product [10 km, hourly]. In validation against measurements from a ground-based sensor network, NEXAI-LST achieves similar or better fit than both LEO and GEO standard products, while depending none of the prior knowledge of land surface and atmospheric states required by physical-statistical models. Further, application of the model to unseen LEO and GEO satellites demonstrates robust generalization of the model across spatial region, time of day, and sensor. In support of NASA’s open-source science initiative, we make our NEXAI-LST product, model, and codes available to facilitate data exploration and further studies.

Kate Marie Duffy↗

A New Machine Learning Based Analysis for Improving Satellite Retrieved Atmospheric Composition Data: OMI SO2 as an Example

Despite recent progress, satellite retrievals of anthropogenic SO2 still suffer from relatively low signal-tonoise ratios. In this study, we demonstrate a new machine learning data analysis method to improve the quality of satellite SO2 products. In the absence of large ground-truth datasets for SO2, we start from SO2 slant column densities (SCDs) retrieved from the Ozone Monitoring Instrument (OMI) using a data-driven, physically based algorithm and calculate the ratio between the SCD and the root mean square (rms) of the fitting residuals for each pixel. To build the training data, we select presumably clean pixels with small SCD / rms ratios (SRRs) and set their target SCDs to zero. For polluted pixels with relatively large SRRs, we set the target to the original retrieved SCDs. We then train neural networks (NNs) to reproduce the target SCDs using predictors including SRRs for individual pixels, solar zenith, viewing zenith and phase angles, scene reflectivity, and O3 column amounts, as well as the monthly mean SRRs. For data analysis, we employ two NNs: (1) one trained daily to produce analyzed SO2 SCDs for polluted pixels each day and (2) the other trained once every month to produce analyzed SCDs for less polluted pixels for the entire month. Test results for 2005 show that our method can significantly reduce noise and artifacts over background regions. Over polluted areas, the monthly mean NN-analyzed and original SCDs generally agree to within ±15 %, indicating that our method can retain SO2 signals in the original retrievals except for large volcanic eruptions. This is further confirmed by running both the NN-analyzed and original SCDs through a topdown emission algorithm to estimate the annual SO2 emissions for ∼ 500 anthropogenic sources, with the two datasets yielding similar results. We also explore two alternative approaches to the NN-based analysis method. In one, we employ a simple linear interpolation model to analyze the original SCD retrievals. In the other, we develop a PCA–NN algorithm that uses OMI measured radiances, transformed and dimension-reduced with a principal component analysis (PCA) technique, as inputs to NNs for SO2 SCD retrievals. While the linear model and the PCA–NN algorithm can reduce retrieval noise, they both underestimate SO2 over polluted areas. Overall, the results presented here demonstrate that our new data analysis method can significantly improve the quality of existing OMI SO2 retrievals. The method can potentially be adapted for other sensors and/or species and enhance the value of satellite data in air quality research and applications.

Can Li↗

A Robust Machine Learning Schema for Developing, Maintaining, and Disseminating Machine Learning Models

Recent advances in the development of machine learning (ML) algorithms have enabled the creation of predictive models that can improve decision making, decrease computational cost, and improve efficiency in a variety of fields. As an organization begins to develop and implement such models, the data used in the training, validation, and testing of ML models, the model parameters, and the use cases or limitations of the models must be properly stored to ensure models are both fully traceable and used correctly. In the context of predicting material behavior, advances in computationally intense, physics-based modeling of material behavior at various length scales and the emergence of Integrated Computational Materials Engineering (ICME) have driven the need for developing data-driven surrogate models of the physics-based simulation tools using ML techniques. Surrogate model development allows for accurate material behavior prediction at a fraction of the cost of its physics-based counterpart, allowing for multiscale simulations of real-world applications, further enabling the ability to design fit-for-purpose materials for a reasonable computational investment. However, training such models requires extensive data, and thus, effective data management is necessary to reach the full potential that ML can offer to material design and ICME. This paper proposes a generalized, robust schema that allows organizations to store both real (experimental) and virtual (simulation) data used to train ML models and the defining model parameters and architectures within the Granta MI Platform. The developed schema allows for various types of data inputs and outputs, including single point values, time-series data, and images that can be used in the prediction of material behavior, while following outlined best practices for effective data management. An effective schema for ML data and models can help prevent the recreation of virtual/real training data and surrogate models, help reduce the time to create new models similar to existing ones by offering a starting point in the hyperparameter determination stages, minimize resources devoted to verification and validation (V&V) and certification of models, and ensure that data and surrogate models are not misused due to full traceability of both the data and ML model. It also allows organizations access to models that have already been developed, such that they can be used in the design of new materials, enabling the overall goals of ICME.

Brandon L. Hearley↗

Validation of Machine Learning Algorithms for Hyperspectral Inversion of Common Water Quality Indicators

The upcoming transition to a diverse suite hyperspectral airborne and orbiting optical sensors will provide an unprecedented opportunity to measure inland water quality characteristics at a fidelity not previously achievable. This presentation will assess prototype deep learning models trained on synthetic hyperspectral data and validated with collocated in-situ measurements. Synthesized data is becoming increasingly popular for use in data-driven approaches to complex problems, and can compliment real data to increase performance on complex and unusual phenomenon, reduce or test bias, and experiment to demonstrate explainability. We will present insights from hyperspectral inversions of Chlorophyl-a, Phycocyanin, and concentration of non-algal particles using selected orbiting and airborne sensors over diverse, optically complex aquatic scenarios. We analyze how various optical water types affect fidelity of results and where improvements can be made as we prototype for globally operational water quality algorithms which can be leveraged by upcoming hyperspectral missions such as the Surface Biology and Geology (SBG) mission.

Surface Biology and Geology (SBG)↗

An Advanced Open-Source Platform for Air Quality Analysis, Visualization, and Prediction

Ambient air pollution is the largest environmental health risk factor, leading to several million premature deaths globally per year. The challenge of combating poor air quality is exacerbated by growing urban populations, changing emissions, and a warming climate. While there have been many advances monitoring and modeling of atmospheric composition, reflected in the dramatic increase in archived Earth Observations, there is no single measurement or method that alone can provide an accurate depiction of the entire atmosphere. The rapidly growing collections of observational and modeling data require us to be smarter about what data to include, and how such data is used. In recent years, NASA has invested significantly in advancing the concepts for Analytics Collaborative Framework (ACF) [5] and New Observing Strategies (NOS) [4] to tackle our software infrastructure need for harmonized data management and dynamic acquisition of diverse measurements for on-demand, interactive, multivariate analysis, and access [3]. It is not enough to have a big data, standalone analytics solution; it is critical that we start integrating data from remote sensing, modeling, and in-situ networks in a harmonized manner that enables timely and data-driven decision-making for air quality management. This work presents the design and development of an Air Quality Analytics Collaborative Framework (AQ ACF), as part of NASA’s Advanced Information Systems Technology (AIST) effort, to establish a data, machine-learning, and numerically driven platform for air quality analysis, visualization, and prediction.

Liu, Qian↗

Prediction of Aircraft Estimated Time of Arrival Using A Supervised Learning Approach

We present a novel data-driven approach for prediction of the estimated time of arrival (ETA) of aircraft in the terminal area via the implementation of a Random Forest regression model. The model uses data fused from a number of sources (flight track, weather, flight plan information, etc.) and provides predictions for the remaining flight time for aircraft landing at Dallas/Fort Worth (DFW) International Airport. The predictions are made when the aircraft is at a distance of 200-miles from the airport. The results show that the model is able to predict estimated time of arrival to within ± 5 min for 90% of the flights in the test data with the mean absolute error being lower at 145 seconds. This paper covers the entire pipeline of data collection, preprocessing, setup and training of the ML model, and the results obtained for DFW.

Machine learning↗

Development of Digital Twin Technologies for Climate Projections

Climate projections are increasingly needed for adaptation, climate resilience and related decision making. However, existing projections have systematic biases, are limited in scope, and are not readily available for most potential users. While the ideal of an observational data-driven ‘digital twin’ for climate is initially attractive, there is only a very limited set of climate data available with which to train such a tool. Nonetheless, we are confident that there is a role for ‘digital twin technologies’ in removing biases, increasing computational efficiency, expanding scenarios and data accessibility.

digital twins↗

FluxSat: Long-term Earth Science Data Record (ESDR) for Terrestrial Gross Primary Production (GPP) based on satellite data calibrated with eddy covariance data

Gross primary production (GPP), the amount of carbon dioxide (CO 2 ) assimilated by plants through photosynthesis, is one of the most variable and uncertain components of the global carbon cycle. Global GPP has been estimated with a number of process-based models, data-driven, and hybrid approaches. Dynamic global vegetation models (DGVMs), driven by observed environmental changes, are used for global carbon budget assessments and long-term (climate) prediction. Benchmarking these and other models globally with data-driven GPP estimates is critical for understanding the land sink and ensuring accurate forecasts of the carbon cycle. In addition, global data-driven GPP estimates are crucial for studies of interannual variability, including trends that are linked to mechanisms with large uncertainties, such as the indirect CO 2 fertilization effect related to greening. In response to a community need for a GPP data set that well captures spatio-temporal variability, we developed FluxSat, a data-driven approach that optimizes the use of satellite reflectance data from the NASA MODerate-resolution Imaging Spectroradiometer (MODIS) on the Terra and Aqua satellites, calibrated using ground-based eddy covariance (EC) data. We are enhancing (spatially, higher resolution) and extending FluxSat (in time, with additional sensors) to create a high quality long term GPP Earth System Data Record (ESDR) for use in model benchmarking, carbon cycle modeling, and studies of trends and interannual variability. Our team’s objectives are to: 1. Update and document the current MODIS FluxSat GPP (daily, 0.05o and 0.5o resolutions) products with latest available MODIS and EC data sets; 2. Extend FluxSat GPP record forward in time with the Visible Infrared Imaging Radiometer Suite (VIIRS) on operational weather satellites going forward; 3. Extend FluxSat GPP record backward in time using the Advanced Very High Resolution Radiometer (AVHRR) on weather satellites dating back to 1981; 4. Provide higher spatial resolution MODIS and VIIRS GPP (0.0083o). 5. Thoroughly evaluate all FluxSat products with independent data; and 6. Create a homogenized long-term GPP record spanning 40+ years. We will discuss plans for this long-term data set that is supported through the NASA Making Earth System Data Records for Use in Research Environments (MEaSUREs) program.

gross Primary Production↗

A Machine Learning Approach to Determine Surface Radiative Fluxes based on CERES Observations

The Clouds and Earth’s Radiant Energy System (CERES) projects provides satellite-based observations of the radiative fluxes and clouds systems. CERES climate quality data products typically take several months of calibration and validation before release to the public. An alternative data product, Fast Longwave and Shortwave radiative Flux (FLASHFlux), was created to provide data to the applied sciences and educational users. FLASHFlux provides Top-of-Atmosphere radiative fluxes, Clouds properties, and parameterized surface radiative fluxes within four days for footprint (Level 2) data. We investigate the use of Artificial Neural Network (ANN) using MODerate resolution Imaging Spectroradiometer (MODIS) derived clouds properties and meteorology from the Global Assimilation and Meteorology Office (GMAO) scaled to the CERES footprint from the CERES Clouds Radiative Swath (CRS) data product to compute surface radiative fluxes. We test ANN produce fluxes against surface fluxes produced from the Fu-Liou model used in CRS and the Langley Parameterized Shortwave Algorithm (LPSA) and Langley Parameterized Longwave Algorithm (LPLA) used in FLASHFlux. We also validated each model with ground-based observations. Furthermore, we investigate Leave-One-Feature-Out Importance (LOFO) to evaluate the significance of each feature in our training and provide insight for future models. Advances in machine learning, along with increases in computational capabilities and available data allow us to estimate effects of unresolved processes in our climate without direct modeling. This work evaluates the ability to create accurate data-driven models to supplement or replace current models that estimate surface radiative fluxes.

Climatology↗

Data-Driven State of Health Estimation for Second-Life Batteries Using Interpolated Synthetic Data and Feature Selection

Accurate estimation of the State of Health (SOH) for second-life batteries (SLBs) is crucial given their increasing use in energy storage applications. Precise SOH prediction is essential for safe operation and robust battery management systems. A major challenge is the limited availability of datasets for building reliable degradation models. To address this, synthetic data generation through linear interpolation is performed to extend the available data, making it more representative of real-world battery operating conditions. By analyzing feature correlation with SOH, the most relevant features are selected for the model. The proposed approach employs a convolutional neural network (CNN) model trained on this interpolated, feature-selected dataset, using time series data of voltage, temperature, and current over a cycle. By focusing on highly correlated features, the model achieves over 95% accuracy, with mean absolute error and root mean squared error up to 2.27% and 2.64%, respectively, in SOH estimation for two battery datasets tested. These results highlight the potential of combining synthetic data generation and feature selection to enhance SOH predictions, showcasing the superior performance of the proposed CNN model for both new batteries and SLBs.

feature selection↗

Predictive Modeling and Uncertainty Quantification in Condition Monitoring of Active Components: A Reactor Coolant Pump Use Case

This work develops data-driven models for onset of thermal barrier leakage in reactor coolant pumps. It incorporates uncertainty quantification to enhance the reliability and robustness of pre- dictions. Using synthetic data generated by the Generic Pressurized Water Reactor simulator, realistic degradation scenarios were simulated across lifecycle stages—beginning, middle, and end of life. Key variables, including differential pressure, flow rate, vibration, and temperatures, were analyzed using machine learning framework. The fully connected neural network models demonstrated exceptional performance, achieving R2 scores exceeding 0.99 and root mean square errors as low as around 8.23 × 10-2 gallon per minute (gpm) for the three stages of the lifecy- cle. UQ analysis further validated the model’s robustness, with narrow uncertainty bounds during steady-state operations and appropriately wider bounds during transitional phases, reflecting the physical behavior of the system. This work addresses important gaps in real-time condition moni- toring and regulatory compliance by integrating advanced condition monitoring technologies with UQ into IST programs. The ability to detect thermal barrier leakage early and quantify prediction reliability supports optimizing maintenance strategies while ensuring nuclear power plants’ safe and reliable operation.

99 - GENERAL AND MISCELLANEOUS↗

Feature-Based PMU Event Classification under Variable PMU Participation and Overlapping Events

This paper is the basis for a presentation help at the 2026 Georgia Tech Fault & Disturbance Analysis Conference, which can be found at OSTI # 3168287 Paper Abstract—Phasor Measurement Units (PMUs) stream time synchronized, high-resolution measurements from the grid, enabling data-driven techniques for event detection and classification. Accurate event classification improves grid reliability and stability. Events can be detected by varying numbers of PMUs and exhibit different durations depending on the event type. This variability challenges standard classifiers that require uniform input sizes. Moreover, multiple events may coincide, which increases classification complexity. Standard classifiers assign each instance to the class with the highest predicted probability, whereas overlapping events may exhibit comparable probabilities across multiple classes. In this study, to handle data size variability, we extract a wide range of time–frequency domain features from all available PMUs for each event into a fixed-length vector, facilitating the application of standard machine learning classifiers, including Random Forest, XGBoost, LightGBM, Support Vector Machine, and Multilayer Perceptron. To account for overlapping events, a probabilistic post-processing step is applied. For a given data instance, if multiple predicted class probabilities exceed 30% and the differences between them are less than 10%, the event is assigned to multiple classes. Experiments using real-world PMU data demonstrate that the Random Forest and XGBoost models achieve the highest accuracy, while the proposed post-processing method yields perfect classification performance on external unseen test sets.

Nematirad, Reza [Danovo Energy Solutions]↗