Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “DATA MINING”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Virtual Sensors: Using Data Mining to Efficiently Estimate Spectra

Detecting clouds within a satellite image is essential for retrieving surface geophysical parameters, such as albedo and temperature, from optical and thermal imagery because the retrieval methods tend to be valid for clear skies only. Thus, routine satellite data processing requires reliable automated cloud detection algorithms that are applicable to many surface types. Unfortunately, cloud detection over snow and ice is difficult due to the lack of spectral contrast between clouds and snow. Snow and clouds are both highly reflective in the visible wavelen,ats and often show little contrast in the thermal Infrared. However, at 1.6 microns, the spectral signatures of snow and clouds differ enough to allow improved snow/ice/cloud discrimination. The recent Terra and Aqua Moderate Resolution Imaging Spectro-Radiometer (MODIS) sensors have a channel (channel 6) at 1.6 microns. Presently the most comprehensive, long-term information on surface albedo and temperature over snow- and ice-covered surfaces comes from the Advanced Very High Resolution Radiometer ( AVHRR) sensor that has been providing imagery since July 1981. The earlier AVHRR sensors (e.g. AVHRR/2) did not however have a channel designed for discriminating clouds from snow, such as the 1.6 micron channel available on the more recent AVHRR/3 or the MODIS sensors. In the absence of the 1.6 micron channel, the AVHRR Polar Pathfinder (APP) product performs cloud detection using a combination of time-series analysis and multispectral threshold tests based on the satellite's measuring channels to produce a cloud mask. The method has been found to work reasonably well over sea ice, but not so well over the ice sheets. Thus, improving the cloud mask in the APP dataset would be extremely helpful toward increasing the accuracy of the albedo and temperature retrievals, as well as extending the time-series of albedo and temperature retrievals from the more recent sensors to the historical ones. In this work, we use data mining methods to construct a model of MODIS channel 6 as a function of other channels that are common to both MODIS and AVHRR. The idea is to use the model to generate the equivalent of MODIS channel 6 for AVHRR as a function of the AVHRR equivalents to MODIS channels. We call this a Virtual Sensor because it predicts unmeasured spectra. The goal is to use this virtual channel 6. to yield a cloud mask superior to what is currently used in APP . Our results show that several data mining methods such as multilayer perceptrons (MLPs), ensemble methods (e.g., bagging), and kernel methods (e.g., support vector machines) generate channel 6 for unseen MODIS images with high accuracy. Because the true channel 6 is not available for AVHRR images, we qualitatively assess the virtual channel 6 for several AVHRR images.

Srivastava, Ashok↗

Data-Mining Toolset Developed for Determining Turbine Engine Part Life Consumption

The current practice in aerospace turbine engine maintenance is to remove components defined as life-limited parts after a fixed time, on the basis of a predetermined number of flight cycles. Under this schedule-based maintenance practice, the worst-case usage scenario is used to determine the usable life of the component. As shown, this practice often requires removing a part before its useful life is fully consumed, thus leading to higher maintenance cost. To address this issue, the NASA Glenn Research Center, in a collaborative effort with Pratt & Whitney, has developed a generic modular toolset that uses data-mining technology to parameterize life usage models for maintenance purposes. The toolset enables a "condition-based" maintenance approach, where parts are removed on the basis of the cumulative history of the severity of operation they have experienced. The toolset uses data-mining technology to tune life-consumption models on the basis of operating and maintenance histories. The flight operating conditions, represented by measured variables within the engine, are correlated with repair records for the engines, generating a relationship between the operating condition of the part and its service life. As shown, with the condition-based maintenance approach, the lifelimited part is in service until its usable life is fully consumed. This approach will lower maintenance costs while maintaining the safety of the propulsion system. The toolset is a modular program that is easily customizable by users. First, appropriate parametric damage accumulation models, which will be functions of engine variables, must be defined. The tool then optimizes the models to match the historical data by computing an effective-cycle metric that reduces the unexplained variability in component life due to each damage mode by accounting for the variability in operational severity. The damage increment due to operating conditions experienced during each flight is used to compute the effective cycles and ultimately the replacement time. Utilities to handle data problems, such as gaps in the flight data records, are included in the toolset. The tool was demonstrated using the first stage, high-pressure turbine blade of the PW4077 engine (Pratt & Whitney, East Hartford, CT). The damage modes considered were thermomechanical fatigue and oxidation/erosion. Each PW4077 engine contains 82 first-stage, high-pressure turbine blades, and data from a fleet of engines were used to tune the life-consumption models. The models took into account not only measured variables within the engine, but also unmeasured variables such as engine health parameters that are affected by degradation of the engine due to aging. The tool proved effective at predicting the average number of blades scrapped over time due to each damage mode, per engine, given the operating history of the engine. The customizable tools are available to interested parties within the aerospace community.

Litt, Jonathan S.↗

A large-scale comparison of Artificial Intelligence and Data Mining (AI&DM) techniques in simulating reservoir releases over the Upper Colorado Region

In recent years, the Artificial Intelligence and Data Mining (AI&DM) models have become popular tools in assisting various aspects of reservoir operation. However, the practical uses are still rarely reported. Comparison experiment of many AI&DM models over a large number of reservoir cases is particularly valuable to help reservoir operators first examine the usefulness and transferability of different AI&DM models, and then identify the most stable and reliable AI&DM model in assist of various decision-making processes. In this study, a total of 12 AI&DM models with different parameterizations and simulation scenarios are comprehensively tested out and compared in simulating the controlled reservoir outflows of 33 reservoir cases over the Upper Colorado Region, United States. Results show that the Random Forecast and the Long-Short-Term-Memory model could consistently derive the best statistical performance than other models under the baseline simulation scenario. The employed AI&DM models could obtain satisfactory statistical interquartile ranges (25–75%) between [0.6–0.9], [0.3–0.8], and [0.2–0.8], for CORR, NSE, and KGE measurements, respectively, and [1.5–6.5], [–15 to 20], and [0.5–8.5] for the normalized RMSE, PBIAS and RSR measurements, respectively. Results also show Multi-Layer Perceptron model and Extreme Gradient Boosting Tree Algorithm produced more stable and superior performance than other models under more complex input scenarios. We also found that the performance of different AI&DM models are closely relevant to the reservoir elevations, sizes, and functionalities. Discussions were made about the sensitivity of AI&DM models’ parameterizations and the key advantages of AI&DM models over the rule-based reservoir models. We further identify that the main advantage of AI&DM models is the flexibility in designing input structures, whereas the rule-based simulation model is rather limited. Future studies were suggested regarding the best way reservoir operators and researchers could use, select, and apply different AI&DM models in simulating reservoir releases under different natural and modeling environments. Finally, this comparison study also serves as a reference and a piece of groundwork for further promoting the practical uses of AI&DM models in assisting reservoir operation.

54 ENVIRONMENTAL SCIENCES↗

Accelerating the discovery of battery electrode materials through data mining and deep learning models

The availability of crystalline materials databases allows for building accurate machine learning (ML) models that can accelerate the exploration of materials chemical space for energy storage applications. In this work, we screen all inorganic materials included in the Materials Project and AFLOW databases as potential metal-ion battery electrodes. We develop an efficient protocol to mine and screen raw data in current databases and provide a new database of electrode materials by considering pairs of charged and discharged electrodes. This effort leads to a new database with over 190,000 instances, in contrast to the original battery database which contains about 5000. The expanded battery data set is then used to build regression-based deep neural network models for predicting average voltages and percentage volume changes upon charging and discharging, which present improvements of at least 28% for target properties with respect to previous models, and are now able to predict anode electrodes (low voltage region) as well as electrodes that will not work in electrochemical cells (negative voltages), overcoming the challenges identified in previous ML models for battery electrodes. Additionally, a further screening of the expanded database itself allowed us to identify 35 novel electrode candidates with excellent battery performance metrics.

25 ENERGY STORAGE↗

Email-Based Informed Consent: Innovative Method for Reaching Large Numbers of Subjects for Data Mining Research

Since the 2010 NASA authorization to make the Life Sciences Data Archive (LSDA) and Lifetime Surveillance of Astronaut Health (LSAH) data archives more accessible by the research and operational communities, demand for data has greatly increased. Correspondingly, both the number and scope of requests have increased, from 142 requests fulfilled in 2011 to 224 in 2014, and with some datasets comprising up to 1 million data points. To meet the demand, the LSAH and LSDA Repositories project was launched, which allows active and retired astronauts to authorize full, partial, or no access to their data for research without individual, study-specific informed consent. A one-on-one personal informed consent briefing is required to fully communicate the implications of the several tiers of consent. Due to the need for personal contact to conduct Repositories consent meetings, the rate of consenting has not kept up with demand for individualized, possibly attributable data. As a result, other methods had to be implemented to allow the release of large datasets, such as release of only de-identified data. However the compilation of large, de-identified data sets places a significant resource burden on LSAH and LSDA and may result in diminished scientific usefulness of the dataset. As a result, LSAH and LSDA worked with the JSC Institutional Review Board Chair, Astronaut Office physicians, and NASA Office of General Counsel personnel to develop a "Remote Consenting" process for retrospective data mining studies. This is particularly useful since the majority of the astronaut cohort is retired from the agency and living outside the Houston area. Originally planned as a method to send informed consent briefing slides and consent forms only by mail, Remote Consenting has evolved into a means to accept crewmember decisions on individual studies via their method of choice: email or paper copy by mail. To date, 100 emails have been sent to request participation in eight HRP-funded studies. The development of the Remote Consent process, the laws allowing transmission of consent via electronic means, total metrics to date, and remaining challenges (e.g., response issues, use of International Partner data, biospecimens/genetic data) for the research use of LSAH/LSDA data will be described.

Lee, Lesley R.↗

Apollo Crew Loads Data Mining

INTRODUCTION Human exposure data to dynamic loads in spaceflight are sparse not only because very few crewmembers have flown in space, but also because much of the historic data is no longer readily available. Earth and Moon landings of the Apollo missions are unique and valuable for understanding human tolerance in water landing conditions for both short-duration and long-duration crewmembers as well as informing crew tolerance in standing postures during planetary landings. The goal of this task was to find relevant Apollo data to better understand human tolerance in these scenarios, in order to apply this knowledge to future missions. METHODS The goal of this task was to recover time history data from dynamic phases of flight from crewed Apollo missions. The first phase of the task was to search available records at JSC and other centers to locate the appropriate records at the National Archives in Fort Worth, TX. Once the possible data locations were located, it was planned that one researcher would travel to the National Archives in Fort Worth, TX to collect the relevant Apollo data. Once the data was retrieved, they would be digitized for future use and evaluation, and archived within the Human Physiology, Performance, Protection, and Operations (H-3PO) laboratory within the Human Health and Performance Directorate. These data would be used to assess the feasibility and acceptability of current standards and vehicle design requirements for lunar and Mars missions. Additionally, any medical monitoring required, or physiological changes noted in crewmembers based on dynamic loads would be documented for future use. RESULTS Relevant records were found at Ft. Worth, TX and College Park, MD using the National Archives Online Catalog. In addition to the online catalog, an Archives staff member was able to share an additional inventory list of NASA JSC records held in Ft. Worth, TX. Due to COVID-19, NASA has only recently allowed non-mission essential travel. In addition, the National Archives has also been closed to the public since March 2020. Because of this, travel to the National Archives to search records there has not been possible. CONCLUSION The findings include collections at the National Archives at both Ft. Worth, TX and College Park, MD that could possibly hold the data we are looking for. Unfortunately, the descriptions of each collection are very broad. Though we are hopeful that relevant data is stored in one of these collections, we cannot be sure that we will find what we are looking for. At this time, it is not clear when the Archives will reopen to the public or when NASA will allow travel. Travel to the Archives has been postponed, and will be reassessed at a future date.

T Reiber↗

A Physics-Informed Reinforcement Learning Framework for Economic-Thermal Co-Optimization of Crypto Mining Data Centers: Preprint

The rapid expansion of cryptocurrency mining has created a new class of high-density data centers characterized by extreme thermal flux and high sensitivity to volatile economic markets. Traditional thermal management strategies, typically reliant on rule-based control, maintain static setpoints that fail to account for fluctuating electricity prices and cryptocurrency values - factors critical to mining profitability. To address this, we present a physics-informed reinforcement learning (PIRL) framework for economic-thermal co-optimization in crypto mining data centers. This framework consists of a proximal policy optimization (PPO) agent, a virtual testbed powered by high-fidelity physics-based models, and an interactive frontend dashboard. The PPO agent is trained using the virtual testbed and strict hardware safety limits. This physics-informed approach allows the agent to learn a stochastic policy that dynamically balances mining revenue against operational costs by co-optimizing HVAC cooling setpoints and IT computational hashrate. The simulation results demonstrate that the integrated framework achieved an 8.62% increase in net operational profit compared to traditional baseline strategies while strictly adhering to safety-critical temperature constraints (coolant supply temperature < 32 degrees C). This work provides a scalable template for the deployment of reinforcement learning in mission critical facilities where economic volatility and physical safety must be managed simultaneously.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Fast Spatio-Temporal Data Mining from Large Geophysical Datasets

Use of the UCLA CONQUEST (CONtent-based Querying in Space and Time) is reviewed for performance of automatic cyclone extraction and detection of spatio-temporal blocking conditions on MPP. CONQUEST is a data analysis environment for knowledge and data mining to aid in high-resolution modeling of climate modeling.

knowledge discovery data mining climate modeling c↗

Accelerated alloy discovery using synthetic data generation and data mining

The search for new alloys with improved properties is never ending with infinite combinations and amounts of alloying elements in the alloy. Advancements in machine learning have made navigating this enormous search space feasible. However, training the machine learning models and tuning their hyper-parameters to make accurate predictions can be time-consuming and often require high-performance computing resources. Furthermore, the quality of the predictions depend on the availability of sufficient training data. Here, we present a generic approach to accelerate alloy discovery by coupling high throughput CALPHAD calculations, synthetic data generation, and data mining. Finally, as a demonstration of the approach, we design super bainitic steels that form bainite at 200 C in lower transformation times.

36 MATERIALS SCIENCE↗

An Exploratory Data Mining Investigation for Constructing a Publicly Sourced Dataset of Foreign Hypersonic Tests

This document details a data mining exercise that resulted in an exploratory dataset of publicly reported foreign (non-US) hypersonic vehicle test events. Using a combination of targeted English language searches and country-specific queries, the study aggregates information from digital news media, official press releases, and social media posts. The resulting list of events captures the publicly available accounts of foreign hypersonic tests, although it does not represent an exhaustive record. Limitations such as inconsistent reporting, translation challenges, and the inherently provisional nature of open-source data are acknowledged. This dataset serves as an initial reference point for further inquiries into high-speed atmospheric phenomena and may facilitate future efforts to correlate these events with geophysical measurements.

33 ADVANCED PROPULSION SYSTEMS↗

An Intelligent Archive Testbed Incorporating Data Mining

Many significant advances have occurred during the last two decades in remote sensing instrumentation, computation, storage, and communication technology. A series of Earth observing satellites have been launched by U.S. and international agencies and have been operating and collecting global data on a regular basis. These advances have created a data rich environment for scientific research and applications. NASA s Earth Observing System (EOS) Data and Information System (EOSDIS) has been operational since August 1994 with support for pre-EOS data. Currently, EOSDIS supports all the EOS missions including Terra (1999), Aqua (2002), ICESat (2002) and Aura (2004). EOSDIS has been effectively capturing, processing and archiving several terabytes of standard data products each day. It has also been distributing these data products at a rate of several terabytes per day to a diverse and globally distributed user community (Ramapriyan et al. 2009). There are other NASA-sponsored data system activities including measurement-based systems such as the Ocean Data Processing System and the Precipitation Processing system, and several projects under the Research, Education and Applications Solutions Network (REASoN), Making Earth Science Data Records for Use in Research Environments (MEaSUREs), and the Advancing Collaborative Connections for Earth-Sun System Science (ACCESS) programs. Together, these activities provide a rich set of resources constituting a value chain for users to obtain data at various levels ranging from raw radiances to interdisciplinary model outputs. The result has been a significant leap in our understanding of the Earth systems that all humans depend on for their enjoyment, livelihood, and survival. The trend in the community today is towards many distributed sets of providers of data and services. Despite this, visions for the future include users being able to locate, fuse and utilize data with location transparency and high degree of interoperability, and being able to convert data to information and usable knowledge in an efficient, convenient manner, aided significantly by automation (Ramapriyan et al. 2004; NASA 2005). We can look upon the distributed provider environment with capabilities to convert data to information and to knowledge as an Intelligent Archive in the Context of a Knowledge Building system (IA-KBS). Some of the key capabilities of an IA-KBS are: Virtual Product Generation, Significant Event Detection, Automated Data Quality Assessment, Large-Scale Data Mining, Dynamic Feedback Loop, and Data Discovery and Efficient Requesting (Ramapriyan et al. 2004).

Ramapriyan, H.↗

Predicting the Operational Acceptance of Airborne Flight Reroute Requests Using Data Mining

For tools that generate more efficient flight routes or reroute advisories, it is important to ensure compatibility of automation and autonomy decisions with human objectives so as to ensure acceptability by the human operators. In this paper, the authors developed a proof of concept predictor of operational acceptability for route changes during a flight. Such a capability could have applications in automation tools that identify more efficient routes around airspace impacted by weather or congestion and that better meet airline preferences. The predictor is based on applying data mining techniques, including logistic regression, a decision tree, a support vector machine, a random forest and Adaptive Boost, to historical flight plan amendment data reported during operations and field experiments. Cross validation was used for model development, while nested cross validation was used to validate the models. The model found to have the best performance in predicting air traffic controller acceptance or rejection of a route change, using the available data from Fort Worth Air Traffic Control Center and its adjacent Centers, was the random forest, with an F-score of 0.77. This result indicates that the operational acceptance of reroute requests does indeed have some level of predictability, and that, with suitable data, models can be trained to predict the operational acceptability of reroute requests. Such models may ultimately be used to inform route selection by decision support tools, contributing to the development of increasingly autonomous systems that are capable of routing aircraft with less human input than is currently the case.

Operational Acceptability↗

Data Mining for ISHM of Liquid Rocket Propulsion Status Update

This document consists of presentation slides that review the current status of data mining to support the work with the Integrated Systems Health Management (ISHM) for the systems associated with Liquid Rocket Propulsion. The aim of this project is to have test stand data from Rocketdyne to design algorithms that will aid in the early detection of impending failures during operation. These methods will be extended and improved for future platforms (i.e., CEV/CLV).

Srivastava, Ashok↗

Data Mining Activities for Bone Discipline - Current Status

The disciplinary goals of the Human Research Program are broadly discussed. There is a critical need to identify gaps in the evidence that would substantiate a skeletal health risk during and after spaceflight missions. As a result, data mining activities will be engaged to gather reviews of medical data and flight analog data and to propose additional measures and specific analyses. Several studies are briefly reviewed which have topics that partially address these gaps in knowledge, including bone strength recovery with recovery of bone mass density, current renal stone formation knowledge, herniated discs, and a review of bed rest studies conducted at Ames Human Research Facility.

Sibonga, J. D.↗

Application of a data-mining method based on Bayesian networks to lesion-deficit analysis

Although lesion-deficit analysis (LDA) has provided extensive information about structure-function associations in the human brain, LDA has suffered from the difficulties inherent to the analysis of spatial data, i.e., there are many more variables than subjects, and data may be difficult to model using standard distributions, such as the normal distribution. We herein describe a Bayesian method for LDA; this method is based on data-mining techniques that employ Bayesian networks to represent structure-function associations. These methods are computationally tractable, and can represent complex, nonlinear structure-function associations. When applied to the evaluation of data obtained from a study of the psychiatric sequelae of traumatic brain injury in children, this method generates a Bayesian network that demonstrates complex, nonlinear associations among lesions in the left caudate, right globus pallidus, right side of the corpus callosum, right caudate, and left thalamus, and subsequent development of attention-deficit hyperactivity disorder, confirming and extending our previous statistical analysis of these data. Furthermore, analysis of simulated data indicates that methods based on Bayesian networks may be more sensitive and specific for detecting associations among categorical variables than methods based on chi-square and Fisher exact statistics.

NASA Discipline Neuroscience↗

Leveraging data mining, active learning, and domain adaptation for efficient discovery of advanced oxygen evolution electrocatalysts

Developing advanced catalysts for acidic oxygen evolution reaction (OER) is crucial for sustainable hydrogen production. This study presents a multistage machine learning (ML) approach to streamline the discovery and optimization of complex multimetallic catalysts. Our method integrates data mining, active learning, and domain adaptation throughout the materials discovery process. Unlike traditional trial-and-error methods, this approach systematically narrows the exploration space using domain knowledge with minimized reliance on subjective intuition. Then, the active learning module efficiently refines element composition and synthesis conditions through iterative experimental feedback. The process culminated in the discovery of a promising Ru-Mn-Ca-Pr oxide catalyst. Our workflow also enhances theoretical simulations with domain adaptation strategy, providing deeper mechanistic insights aligned with experimental findings. By leveraging diverse data sources and multiple ML strategies, we demonstrate an efficient pathway for electrocatalyst discovery and optimization. This comprehensive, data-driven approach represents a paradigm shift and potentially benchmark in electrocatalysts research.

Science & Technology - Other Topics↗

Data Mining for Anomaly Detection

The Vehicle Integrated Prognostics Reasoner (VIPR) program describes methods for enhanced diagnostics as well as a prognostic extension to current state of art Aircraft Diagnostic and Maintenance System (ADMS). VIPR introduced a new anomaly detection function for discovering previously undetected and undocumented situations, where there are clear deviations from nominal behavior. Once a baseline (nominal model of operations) is established, the detection and analysis is split between on-aircraft outlier generation and off-aircraft expert analysis to characterize and classify events that may not have been anticipated by individual system providers. Offline expert analysis is supported by data curation and data mining algorithms that can be applied in the contexts of supervised learning methods and unsupervised learning. In this report, we discuss efficient methods to implement the Kolmogorov complexity measure using compression algorithms, and run a systematic empirical analysis to determine the best compression measure. Our experiments established that the combination of the DZIP compression algorithm and CiDM distance measure provides the best results for capturing relevant properties of time series data encountered in aircraft operations. This combination was used as the basis for developing an unsupervised learning algorithm to define "nominal" flight segments using historical flight segments.

Biswas, Gautam↗