Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Machine learning prediction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

LEGION: Lightweight Expandable Group of Independently Operating Nodes

LEGION is a lightweight C-language software library that enables distributed asynchronous data processing with a loosely coupled set of compute nodes. Loosely coupled means that a node can offer itself in service to a larger task at any time and can withdraw itself from service at any time, provided it is not actively engaged in an assignment. The main program, i.e., the one attempting to solve the larger task, does not need to know up front which nodes will be available, how many nodes will be available, or at what times the nodes will be available, which is normally the case in a "volunteer computing" framework. The LEGION software accomplishes its goals by providing message-based, inter-process communication similar to MPI (message passing interface), but without the tight coupling requirements. The software is lightweight and easy to install as it is written in standard C with no exotic library dependencies. LEGION has been demonstrated in a challenging planetary science application in which a machine learning system is used in closed-loop fashion to efficiently explore the input parameter space of a complex numerical simulation. The machine learning system decides which jobs to run through the simulator; then, through LEGION calls, the system farms those jobs out to a collection of compute nodes, retrieves the job results as they become available, and updates a predictive model of how the simulator maps inputs to outputs. The machine learning system decides which new set of jobs would be most informative to run given the results so far; this basic loop is repeated until sufficient insight into the physical system modeled by the simulator is obtained.

Burl, Michael C.↗

Machine Learning to Increase the Quality and Repeatability of 3D Printing - Workflow

The imprecise nature of three-dimensional (3D) printing limits the technology’s use beyond prototyping. For production of end-use parts, such as those for aerospace applications, improvements are needed to enhance quality and repeatability. Much of the difficulty in obtaining high quality printed parts lies in finding optimum printing parameters. Currently, this requires trial and error performed by an expert. Finding the optimum printing parameters is also obfuscated by the variation in optimum parameters throughout the part due to part geometry and printer effects. To allow for locally optimized printing parameters, one can envision a machine learning algorithm that could take in an object, predict the best printing parameters, and communicate these parameters to a printer. With this scenario in mind, we developed a tool that can predict and implement locally optimized printing parameters in 3D printing. This tool consists of elements designed to detect errors in a printed part, predict the probability of local flaws occurring at each point in the part, and select the optimal local parameters for the highest quality part given hardware limitations. The results of this work were highlighted in Advanced Materials Technologies. In this paper, we will discuss in greater depth the workflow and algorithms involved with this tool that were not detailed in the journal publication.

additive manufacturing↗

Flight Trajectory Prediction Based on Hybrid-Recurrent Networks

The development of future technologies for the National Airspace System (NAS) will be reliant on a new communications infrastructure capable of managing the limited available spectrum for communications among aircraft and ground systems. Emerging approaches to autonomous allocation of aviation spectrum mostlyrely on machine learning techniques, where 4D (longitude, latitude, altitude, time) trajectory prediction is an important data input to enable real-time resource allocation. This study explores and evaluates effective data sources and deep recurrent neural network techniques when determining flight trajectories. Specifically, data are collected and evaluated in a 100-day and 14-day period. Sources of data include NASA Sherlock Data Warehouse, MIT Lincoln Labs Corridor Integrated Weather Service (CIWS), and assorted NOAA weather datasets. Deep learning models for 4D predictions all utilize a hybrid-recurrent technique. A baseline model is considered via the convolutional-LSTM design from the existing literature. The modified design considers Gated Recurrent Units (GRU), Independently Recurrent Neural Networks (IndRNN), and stand-alone self-attention layers. Results indicatethe effectiveness of LSTM and GRUcells for state-of-the-art data processing (interpolation). Additionally, GRUs may be quickly trained with limited data, allowing for exacting improvements with optimizer selection. Attention mechanisms provide notable performance improvements to convolutional layers and may extend dimensional capabilities of a learning model. Finally, NOAA measurements provide only a supplemental value, requiring support from tailored measurements for Air Traffic Management.

Nathan Schimpf↗

MARGInS Model-Based Analysis of Realizable Goals in Systems

The high complexity of modern aircraft and spacecraft requires elaborate Verification and Validation (V&V) approaches to make sure that such complex systems work properly and reliably. MARGInS is a framework for the analysis, understanding, and prediction of the behavior of a complex, hybrid system. MARGInS contains a set of machine learning and statistical algorithms for multivariate clustering, treatment learning, critical factor determination, time-series analysis, event prediction, and safety-boundary detection and characterization. The framework supports system testing and can be configured to find novel features in test suites, determine classes of behavior, propose new experiments that can efficiently explore and characterize the boundaries between classes of system behavior, and to create visualizations and reports.

He, Yuning↗

EARLY INFORMATION PARAMETER-SET ANALYSIS FOR SATELLITE CLOSE APPROACHES USING MACHINE LEARNING

In spaceflight navigation applications, understanding and accurately applying orbital mechanics by leveraging force models for trajectory predictions will always remain an important aspect in space mission design and operations. In the process of capturing the dynamics and perturbations in the space environment, the force models are not all encompassing in that these models are subject to errors, commonly referred to as process noise. Therefore, in predicting state vectors of space objects such as spacecraft or debris over long periods of time, these errors in the process noise tend to grow over time.

machine learning↗

A Robust Machine Learning Schema for Developing, Maintaining, and Disseminating Machine Learning Models

Recent advances in the development of machine learning (ML) algorithms have enabled the creation of predictive models that can improve decision making, decrease computational cost, and improve efficiency in a variety of fields. As an organization begins to develop and implement such models, the data used in the training, validation, and testing of ML models, the model parameters, and the use cases or limitations of the models must be properly stored to ensure models are both fully traceable and used correctly. In the context of predicting material behavior, advances in computationally intense, physics-based modeling of material behavior at various length scales and the emergence of Integrated Computational Materials Engineering (ICME) have driven the need for developing data-driven surrogate models of the physics-based simulation tools using ML techniques. Surrogate model development allows for accurate material behavior prediction at a fraction of the cost of its physics-based counterpart, allowing for multiscale simulations of real-world applications, further enabling the ability to design fit-for-purpose materials for a reasonable computational investment. However, training such models requires extensive data, and thus, effective data management is necessary to reach the full potential that ML can offer to material design and ICME. This paper proposes a generalized, robust schema that allows organizations to store both real (experimental) and virtual (simulation) data used to train ML models and the defining model parameters and architectures within the Granta MI Platform. The developed schema allows for various types of data inputs and outputs, including single point values, time-series data, and images that can be used in the prediction of material behavior, while following outlined best practices for effective data management. An effective schema for ML data and models can help prevent the recreation of virtual/real training data and surrogate models, help reduce the time to create new models similar to existing ones by offering a starting point in the hyperparameter determination stages, minimize resources devoted to verification and validation (V&V) and certification of models, and ensure that data and surrogate models are not misused due to full traceability of both the data and ML model. It also allows organizations access to models that have already been developed, such that they can be used in the design of new materials, enabling the overall goals of ICME.

Brandon L. Hearley↗

Bryce Canyon Water Resources: Monitoring Vegetation Health and Water Availability in Bryce Canyon National Park for Drought Stress Mitigation Planning

Bryce Canyon National Park is home to groundwater-dependent ecosystems (GDEs) that are threatened by a multidecadal drought and increased groundwater extraction due to a spike in tourism. These ecosystems contain unique species that are only found in areas where near-surface groundwater is present, such as aspen groves and fens. These species contribute to the high biodiversity found in Bryce Canyon, which boosts an ecosystem’s productivity and the services it provides to the park. Unfortunately, many of these GDEs are too small to identify with traditional Earth observation platforms and are difficult to physically reach for monitoring purposes. This project partnered with the National Park Service to identify springs and seeps as a proxy for GDEs within Bryce Canyon from 2013–2022. Furthermore, this project tested the feasibility of various methods to detect and monitor springs and seeps and therefore facilitate the partner’s efforts to conserve these ecologically valuable GDEs in Bryce Canyon. The team mapped groundwater discharge with high resolution National Agriculture Imagery Program (NAIP) and assessed park vegetation trends with Landsat 8 Operational Land Imager (OLI) and PlanetScope imagery. In-situ precipitation data and the Western Land Data Assimilation System (WLDAS) were used to produce time series of climatic variables. Seeps and spring locations were predicted using random forest classification and maximum entropy machine learning models.

Groundwater dependent ecosystems↗

Information Systems Technology for NASA Earth Systems Digital Twins (ESDT)

The term “Digital Twin” was first used in 2002 for product lifecycle management. Since then, Digital Twin concepts have been proposed in various domains until very recently for Earth Science. For NASA’s Advanced Information Systems Technology (AIST) Program, an Earth System Digital Twin (ESDT) is defined as composed of three components: 1. A Digital Replica, i.e., an integrated picture of the past and current states of Earth systems 2. Forecasting capabilities, providing an integrated picture of how Earth systems will evolve in the future from the current state 3. Impact Assessment capabilities, providing an integrated picture of how Earth systems could evolve under different hypothetical what-if scenarios. Developing such a vision will require technologies related to: integrating continuous observations from various disparate sources; developing frameworks that builds on inter-connected models; improving the speed and accuracy of integrated prediction, analysis and visualization capabilities (e.g., by using machine learning); and utilizing causality and uncertainty quantification to improve our understanding of the evolution of Earth Science systems as a function of their interactions with other Earth and human systems. In addition, AIST is also investigating interoperability standards to federate multiple Digital Twins, as well as computational resources required by those systems.

Earth Science Remote Sensing; Information Systems↗

Machine Learning Algorithms for Alignment Verification of the Roman Space Telescope

The Nancy Grace Roman Telescope is a NASA observatory designed to unravel the secrets of dark energy and dark matter, search for and image exoplanets, and explore many topics in infrared optics. Scheduled to launch no earlier than October 2026, this 2.4 meter aperture telescope has a field of view 100 times greater than the Hubble Space Telescope. The mission is currently in its construction phase, where the telescope and its two instruments will soon be aligned together to ensure proper pupil matching. To help verify this alignment, multiple point sources above the entrance pupil of the telescope will illuminate the optical path through the telescope-instrument system, and shadows of various obstructions in the system will be analyzed using machine learning algorithms to determine the pupil matching error. This presentation discusses the test approach and the machine learning algorithms employed, as well as our uncertainty predictions based on a modeled Monte-Carlo analysis of the test.

Telescope↗

Predictive Modeling for Differential Diagnosis and Mortality Risk Assessment

The prevalence of electronic health record (EHR) systems has brought prodigious biomedical informatics opportunity. Automated machine learning methods can effectively utilize such data and have become common tools for healthcare predictive modeling. Researches in medical informatics have explored the potential of deep learning and classical models in emergent care scenarios. In particular, predicting differential diagnoses for admissions have proven useful in decreasing unnecessary lab tests and improving inpatient triage decision-making. Moreover, identification of high-risk patients for in-hospital mortality is vitally important to maximize allocation of medical resources.The Medical Information Mart for Intensive Care (MIMIC-III) database, containing de-identified critical care inpatient was used in our study. This data set captures hospital patient laboratory measurements, pharmacologic prescriptions, diagnostic data and procedure event recordings. When considering adult patients and discounting admissions with ICU length of stay less than 24 hours, there were 37,787 unique admissions and 30,414 total patients. We examined the top 25 most prevalent ICD-9 group-level disease specificities in MIMIC-III using a multi-label classification model. In-hospital mortality was modeled as binary classification with 4,155 (13%) adult patients that expired, of which 3,138 (75.5%) were in the ICU setting. The metrics AUC, F1 score, sensitivity and specificity values calculated for each disease label measured prediction performance.The usage of ICD-9 group codes reduced feature dimension from 14,567 to 942 and greatly improved distribution of patient diagnostic categories. Disease temporal patterns were captured by considering the most frequently sampled 6 vital signs and 13 laboratory values. Missing data were imputed at each time-stamp. Time-series raw hourly average values were converted into 5 summary features (mean, standard deviation, number of observations, min & max values). Patient demographic variables such as age, gender, marital status and ethnicity were also factored into the modeling. Choi et al showed that contextual embedding of medical data, diagnostic and procedural codes alone can predict future diagnoses with sensitivity as high as 0.79. We utilized an embedding technique called word2vec which allowed sparse representations of medical history to be transformed into dense word vectors. The mappings captured contextual information by treating each admission as a sentence and learning the most likely neighboring words in a sliding window fashion. Binary and multi-label classification was achieved via collapse models, which do not consider temporal information, as well as recurrent neural networks with regularization, Softmax output layer activation together with categorical cross-entropy as the loss function.

US Army collaboration↗

Applications of Principled Search Methods in Climate Influences and Mechanisms

Forest and grass fires cause economic losses in the billions of dollars in the U.S. alone. In addition, boreal forests constitute a large carbon store; it has been estimated that, were no burning to occur, an additional 7 gigatons of carbon would be sequestered in boreal soils each century. Effective wildfire suppression requires anticipation of locales and times for which wildfire is most probable, preferably with a two to four week forecast, so that limited resources can be efficiently deployed. The United States Forest Service (USFS), and other experts and agencies have developed several measures of fire risk combining physical principles and expert judgment, and have used them in automated procedures for forecasting fire risk. Forecasting accuracies for some fire risk indices in combination with climate and other variables have been estimated for specific locations, with the value of fire risk index variables assessed by their statistical significance in regressions. In other cases, the MAPSS forecasts [23, 241 for example, forecasting accuracy has been estimated only by simulated data. We describe alternative forecasting methods that predict fire probability by locale and time using statistical or machine learning procedures trained on historical data, and we give comparative assessments of their forecasting accuracy for one fire season year, April- October, 2003, for all U.S. Forest Service lands. Aside from providing an accuracy baseline for other forecasting methods, the results illustrate the interdependence between the statistical significance of prediction variables and the forecasting method used.

Glymour, Clark↗

Interpretable Machine Learning for Molecular Biosignatures: a Novel Single-Sample Feature Importance Method That Is Sensitive To Statistical Interactions

Isotope ratio mass spectrometry (IRMS) of volatiles (e.g., CO 2 ) promises to be a powerful tool for potential biosignature detection for future missions to ocean worlds (OW) such as Europa and Enceladus. Machine learning (ML) methods for IRMS data could enable science autonomy by onboard prediction of seawater chemistry and biosignature presence. However, ML models are likely to be complex and involve statistical interactions between features (variables), which can make predictions seem opaque and enigmatic. For ML predictions as significant as extraterrestrial biosignatures, we must place extraordinary confidence in models. It is therefore essential that these models make interpretable predictions (i.e., human-understandable) and include false-prediction diagnostics. We achieve high accuracy and interpretability in ML biosignature and seawater chemistry models for OW through a nearest-neighbors feature selection tool that detects statistical interactions between predictors, constructs interaction networks for visualization of selected features working together to make a prediction, and reports single-sample feature importance scores for false-detection diagnostics. Here we develop a novel single-sample nearest-neighbors projected distance regression(ssNPDR) feature selection method that improves upon existing single-sample algorithms through the inclusion of statistical interactions while providing false-prediction diagnostics for ML models.

geochemistry↗

Using Machine Learning for Timely Estimates of Ocean Color Information From Hyperspectral Satellite Measurements in the Presence of Clouds, Aerosols, and Sunglint

Retrievals of ocean color from space are important for better understanding of the ocean ecosystem but can be limited under conditions such as clouds, aerosols, and sunglint. Many ocean color algorithms use a few selected spectral bands to perform an atmospheric correction and then derive the upwelling radiance from the ocean. The limitations in the atmospheric correction under certain conditions lead to many gaps in daily spatial coverage of ocean color retrievals. To address these limitations, we introduce a new approach that uses machine learning to estimate ocean color from top of atmosphere radiances or reflectance measurements. In this approach, a principal component analysis is used to decompose the hyperspectral measurements into spectral features that describe the scattering and absorption of the atmosphere and the underlying surface. The coefficients of the principal components are then used to train a neural network to predict ocean color properties derived from the MODIS atmospheric correction algorithm. This machine learning approach is independent of a priori information and does not rely on any radiative transfer modeling. We apply the approach to two hyperspectral UV/VIS instruments, the ozone monitoring instrument (OMI) and the TROPOspheric Monitoring Instrument (TROPOMI), using measurements from 320–500 nm to show that it can be used to reproduce ocean color properties in less-than-ideal conditions. This machine learning approach complements the current atmospheric correction ocean color retrievals by filling in the gaps resulting from cloud, aerosol, and sunglint contamination. This method can be applied to the future hyperspectral Ocean Color Instrument (OCI), which will be onboard NASA’s Plankton, Aerosol Cloud, ocean Ecosystem (PACE) ocean color satellite set to launch in 2024.

Ocean color↗

Digitally-Engineered Impact Resistant Aerogel Composites for MMOD Protection (DIRAC-MP)

This project implemented a digital-engineering approach to optimize the impact absorption performance of polymer aerogels and aerogel-based composites for Micrometeoroids and Orbital Debris (MMOD) containment. We developed a curated materials database and a machine-learning framework to derive composition-response relationships, enabling predictive design and targeted material selection. In support of experimental validation, a split Hopkinson pressure bar (SHPB) test rig, specifically adapted for low-density aerogel materials, was designed and built in-house. This project accelerates the development of new aerogel formulations, producing candidate materials tailored for enhanced impact-absorption behavior.

Sadeq Malakooti↗

Use of TEMPO as a Proxy for Hyperspectral Geostationary Ocean Color Measurements from the GeoXO OCX Instrument: Harnessing Machine Learning and Principal Component Techniques for Atmospheric and Glint Correction

Retrievals of ocean color from space are important for better understanding the ocean ecosystem. The launch of atmospheric geostationary hyperspectral sensors such as TEMPO, provides a unique opportunity to examine the diurnal variability in ocean ecology. While TEMPO does not have as high spatial resolution or full spectral coverage as planned coastal ocean sensors such as the Geosynchronous Littoral Imaging and Monitoring Radiometer (GLIMR) or GeoXO Ocean Color instrument (OCX), its hourly measurements provide coverage of regions such as Lake Erie and the Gulf of Mexico at spatial scales of approximately 5 km. These data can be useful for testing new algorithms. We will apply our newly developed machine learning based atmospheric correction approach for ocean color retrievals to TEMPO data. Our approach begins by decomposing measured radiances from hyperspectral sensors into spectral features that describe the scattering and absorption of the atmosphere as well as the underlying surface reflectance. The coefficients of the principal components are then used to train a neural network to predict ocean color properties derived from collocated MODIS/VIIRS physically-based retrievals. This machine learning approach does not rely on radiative transfer modeling, and the use of MODIS/VIIRS data for training accounts for possible calibration b in hyperspectral data. Previously, we applied our approach using blue and UV wavelengths with the Ozone Monitoring Instrument (OMI) and TROPOspheric Monitoring Instrument (TROPOMI) to show that it can estimate ocean color properties in less-than-ideal conditions such as lightly to moderately clouded conditions as well as sun glint and thus improve the spatial coverage of ocean color measurements. TEMPO provides an opportunity to improve on this approach since it will provide collocated measurements at green and red wavelengths that were not available from OMI and TROPOMI and are important particularly for coastal waters. Additionally, our technique can be applied early in the mission and has potential to demonstrate the value of near real time ocean color products that are important for monitoring of harmful algae blooms and other oceanic phenomena.

Zachary Fasnacht↗

Predicting Air Traffic Management Initiatives Using Supervised Learning

Terminal Traffic Management Initiatives (TMIs) such as Ground Stops (GS) and Ground Delay Programs (GDP) are implemented to manage excess demand or lowered capacity at an airport. Air Traffic Flow Management (TFM) specialists identify situations such as aviation constraints, current and forecasted weather conditions, airport demand and capacity, and initiate TMIs for safe and orderly movement of air traffic. In this paper, we outline supervised learning techniques that can be used to predict and recommend TMIs at an airport based on current weather and airport conditions. Our research involves building classic Machine Learning (ML) models such as Logistic Regression, K-Nearest Neighbor, Random Forest and XGBoost, as well as Long short-term memory (LSTM) networks. We trained the models on 3-year historical data (weather, airport demand, capacity and TMIs) from Newark (EWR) airport which was selected based on its higher TMI implementation rates and varied weather conditions. Although Random Forest and XGBoost algorithms are able to predict if a TMI is needed or not, they have difficulty in predicting specific program type. For this purpose, we found that LSTM time-series forecasting models performed better as they also learn from past TMI program type sequences. This study also lays down the foundation for advanced modeling techniques and architectures to predict TMIs in advance for future periods. The ability to predict TMIs in advance will be highly beneficial to the traffic controllers and managers as this will help them to prepare for and manage TMIs more efficiently.

Manoj Agrawal↗

Modeling Key Predictors of Airport Runway Configurations Using Learning Algorithms

Advanced traffic flow management automation will need accurate predictions of airport runway configurations. Terminal area weather and traffic demand are generally considered to be the most significant factors in predicting runway configuration. Weather information is forecasted across multiple features, including wind direction, wind speed, gusts, cloud ceilings, visibility, temperature, and precipitation, among many others. We use machine learning techniques on historical weather and runway data to determine weather features that correlate well with runway configurations. We analyze the predictive capability of weather features using different learning models trained on data from four major U.S. airports: Atlanta (ATL), Washington – Dulles (IAD), New York – Kennedy (JFK), and San Francisco (SFO). Wind direction alone is strongly correlated with runway configurations above all other examined factors, as expected. This correlation is the most significant component of the ~80% prediction accuracy in selecting between the two most frequently used runway configurations. However, individual airports show variations on how well the runway configuration decisions correlate with wind direction. While wind direction was identified as the most significant indicator of configuration decisions in ATL, IAD, and JFK, it did not emerge as such at SFO. Traffic demand was not found to be a strong factor in predicting runway configurations at any of the airports analyzed. In rare instances, when high demand cannot be accommodated within the current configuration, temporary changes are likely to be attributable to demand. However, these occurrences are so limited in number that their overall effect is not sufficient to consider traffic demand as a major indicator of runway configuration at the airports analyzed.

Bilimoria, Karl D.↗

Identification of Flux Rope Orientation via Neural Networks

Geomagnetic disturbance forecasting is based on the identification of solar wind structures and accurate determination of their magnetic field orientation. For nowcasting activities, this is currently a tedious and manual process. Focusing on the main driver of geomagnetic disturbances, the twisted internal magnetic field of interplanetary coronal mass ejections (ICMEs), we explore a convolutional neural network’s (CNN) ability to predict the embedded magnetic flux rope’s orientation once it has been identified from in situ solar wind observations. Our work uses CNNs trained with magnetic field vectors from analytical flux rope data. The simulated flux ropes span many possible spacecraft trajectories and flux rope orientations. We train CNNs first with full duration flux ropes and then again with partial duration flux ropes. The former provides us with a baseline of how well CNNs can predict flux rope orientation while the latter provides insights into real-time forecasting by exploring how accuracy is affected by percentage of flux rope observed. The process of casting the physics problem as a machine learning problem is discussed as well as the impacts of different factors on prediction accuracy such as flux rope fluctuations and different neural network topologies. Finally, results from evaluating the trained network against observed ICMEs from Wind during 1995–2015 are presented.

Thomas Narock↗