Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “DATA MINING”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Data Mining for Anomaly Detection

The Vehicle Integrated Prognostics Reasoner (VIPR) program describes methods for enhanced diagnostics as well as a prognostic extension to current state of art Aircraft Diagnostic and Maintenance System (ADMS). VIPR introduced a new anomaly detection function for discovering previously undetected and undocumented situations, where there are clear deviations from nominal behavior. Once a baseline (nominal model of operations) is established, the detection and analysis is split between on-aircraft outlier generation and off-aircraft expert analysis to characterize and classify events that may not have been anticipated by individual system providers. Offline expert analysis is supported by data curation and data mining algorithms that can be applied in the contexts of supervised learning methods and unsupervised learning. In this report, we discuss efficient methods to implement the Kolmogorov complexity measure using compression algorithms, and run a systematic empirical analysis to determine the best compression measure. Our experiments established that the combination of the DZIP compression algorithm and CiDM distance measure provides the best results for capturing relevant properties of time series data encountered in aircraft operations. This combination was used as the basis for developing an unsupervised learning algorithm to define "nominal" flight segments using historical flight segments.

Biswas, Gautam↗

Reports From EXPORTS Modeling and Data-Mining Activities

EXPORTS -EXport Processes in the Ocean from RemoTe Sensing is NASA’s large field campaign focusing on development of a predictive understanding of the export, fate and carbon cycle impacts of the global net primary production. This co-funded program (NSF, private funding1, and international participation) was conceived in 2013, with EXPORTS Science plan published in 2016 (Siegel et al 2016, EXPORTS Writing Team 2015) and its implementation plan finalized in 2016 (EXPORTS Science Definition Team 2016). More details are presented in Siegel et al (2021). EXPORTS campaign is structured as a multiyear effort (Figure 1). It started with “Pre-EXPORTS” modeling and data-mining activity followed by a first phase with two major field programs and a second synthesis and modeling phase. The “Pre-EXPORTS” projects, total of 6 of them (Table 1), funded under A.3 Ocean Biology and Biogeochemistry 2015 call, helped to plan the field campaign (Resplandy et al 2019, Rousseaux & Gregg 2017), and supporting further global synthesis with datasets mined from the literature (e.g. Bisson et al (2020), Bisson et al (2018), Kramer and Siegel (2019)), directly responding to objectives outlined in Science Plan (see section 6 in Team (2015)) and Implementation team (see Figure 1 in Team (2016)). This document presents a compilation of the final reports of the Pre-EXPORTS funded projects, in hope of synthesizing the outcomes, and insuring the legacy of this program. Each of these reports contains a list of published papers, and reader should refer to them to see results in details.

Brandi J. McCarty↗

Data Mining of Polymer Phase Transitions upon Temperature Changes by Small and Wide-Angle X-ray Scattering Combined with Raman Spectroscopy

The complex physical transformations of polymers upon external thermodynamic changes are related to the molecular length of the polymer and its associated multifaceted energetic balance. The understanding of subtle transitions or multistep phase transformation requires real-time phenomenological studies using a multi-technique approach that covers several length-scales and chemical states. A combination of X-ray scattering techniques with Raman spectroscopy and Differential Scanning Calorimetry was conducted to correlate the structural changes from the conformational chain to the polymer crystal and mesoscale organization. Current research applications and the experimental combination of Raman spectroscopy with simultaneous SAXS/WAXS measurements coupled to a DSC is discussed. In particular, we show that in order to obtain the maximum benefit from simultaneously obtained high-quality data sets from different techniques, one should look beyond traditional analysis techniques and instead apply multivariate analysis. Data mining strategies can be applied to develop methods to control polymer processing in an industrial context. Crystallization studies of a PVDF blend with a fluoroelastomer, known to feature complex phase transitions, were used to validate the combined approach and further analyzed by MVA.

36 MATERIALS SCIENCE↗

High Performance EVA Glove Collaboration: Glove Injury Data Mining Effort

Human hands play a significant role during Extravehicular Activity (EVA) missions and Neutral Buoyancy Lab (NBL) training events, as they are needed for translating and performing tasks in the weightless environment. Because of this high frequency usage, hand and arm related injuries are known to occur during EVA and EVA training in the NBL. The primary objectives of this investigation were to: 1) document all known EVA glove related injuries and circumstances of these incidents, 2) determine likely risk factors, and 3) recommend interventions where possible that could be implemented in the current and future glove designs. METHODS: The investigation focused on the discomforts and injuries of U.S. crewmembers who had worn the pressurized Extravehicular Mobility Unit (EMU) spacesuit and experienced 4000 Series or Phase VI glove related incidents during 1981 to 2010 for either EVA ground training or in-orbit flight. We conducted an observational retrospective case-control investigation using 1) a literature review of known injuries, 2) data mining of crew injury, glove sizing, and hand anthropometry databases, 3) descriptive statistical analyses, and finally 4) statistical risk correlation and predictor analyses to better understand injury prevalence and potential causation. Specific predictor statistical analyses included use of principal component analyses (PCA), multiple logistic regression, and survival analyses (Cox proportional hazards regression). Results of these analyses were computed risk variables in the forms of odds ratios (likelihood of an injury occurring given the magnitude of a risk variable) and hazard ratios (likelihood of time to injury occurrence). Due to the exploratory nature of this investigation, we selected predictor variables significant at p≤0.15. RESULTS: Through 2010, there have been a total of 330 NASA crewmembers, from which 96 crewmembers performed 322 EVAs during 1981-2010, resulting in 50 crewmembers being injured inflight and 44 injured during 11,704 ground EVA training events. Of the 196 glove related injury incidents, 106 related to EVA and 90 to EVA training. Over these 196 incidents, 277 total injuries (126 flight; 151 training) were reported and were then grouped into 23 types of injuries. Of EVA flight injuries, 65% were commonly reported to the hand (in general), metacarpophalangeal (MCP) joint, and finger (not including thumb) with fatigue, abrasion, and paresthesia being the most common injury types (44% of total flight injuries). Training injuries totaled to more than 70% being distributed to the fingernail, MCP joint, and finger crotch with 88% of the specific injuries listed as pain, erythema, and onycholysis. Of these training injuries, when reporting pain or erythema, the most common location was the index finger, but when reporting onycholysis, it was the middle finger. Predictor variables specific to increased risk of onycholysis included: female sex (OR=2.622), older age (OR=1.065), increased duration in hours of the flight or training event (OR=1.570), middle finger length differences in inches between the finger and the EVA glove (OR=7.709), and use of the Phase VI glove (OR=8.535). Differentiation between training and flight and injury reporting during 2002-2004 were significant control variables. For likelihood of time to first onycholysis injury, there was a 24% reduction in rate of reporting for each year increase in age. Also, more experienced crewmembers, based on number of EVA flight or training events completed, were less likely to report an onycholysis injury (3% less for every event). Longer duration events also found reporting rates to occur 2.37 times faster for every hour of length. Crewmembers with larger hand size reported onycholysis 23% faster than those with smaller hand size. Finally, for every 1/10th of an inch increase in difference between the middle finger length and the glove, the rate of reporting increased by 60%. DISCUSSION: One key finding was that the Series 4000 glove had a lower injury risk than the Phase VI, which provides a platform for further evaluation. General interventions that reduce hand overexertion and repetitive use exposure through tool development, procedural changes and shorter exposures may be one mitigation path, but due to the way the training event times were reported, we cannot provide a guideline for a specific event duration change. When the finger length was different from the glove length, the risk of injury increased indicating that the use of larger finger take-ups could be contributing to injury and therefore may not be recommended. Prior to this investigation, there was one previous investigation indicating hand anthropometry may be related to onycholysis. We found different hand anthropometry variables indicated by this investigation as compared to the prior, specifically differences in middle finger length compared to glove finger length, which point more towards a sizing issue than a specific anthropometry issue. Additionally, although this investigation has identified sizing as an issue, the force and environmental-related variables of the EVA glove that could also cause injury were not accounted for.

Reid, C. R.↗

Data Mining and Machine Learning for Power System Monitoring, Understanding, and Impact Evaluation

This chapter presents results from the Big Data analysis framework to improve power system situational awareness and system reliability. For this purpose, a dataset with real-world phasor measurement unit data and historical transmission system outage data has been created and used to carry out the analysis. Several statistical analysis and machine learning methods have been developed and implemented for event and anomaly detection and modeling. Detection and analysis results for actual examples of power system events are presented. Finally, data-driven characterization and risk assessment methods for weather-related extremes in power systems are developed and demonstrated on the Bonneville Power Administration system. These applications demonstrate the capability of Machine Learning (ML) methods to monitor system abnormalities, to predict system events, and to characterize the impact of extreme events on power grid

data mining, power grid, machine learning, anomaly↗

Data mining of plug-in electric vehicles charging behavior using supply-side data

This paper aims to better understand the charging patterns of plug-in electric vehicles (PEVs) and identify factors that may significantly impact PEVs’ charging behavior. We collected 189,864 supply-side charging session data over 13 months from 821 charging stations in Illinois from ChargePoint. Through descriptive and regression analyses, we characterize the distributions of key charging behavior indicators, including charging location, dwell time, and battery start state of charge (SOC), and quantify the impacts of closely related factors on these charging behaviors. In this work, we find that: (1) PEVs are more likely to charge in the morning at multifamily commercial locations with a lower start SOC compared with single family residential locations; (2) Weekday and morning sessions are more likely to utilize workplace charging and have shorter dwell time compared with weekend and afternoon sessions; (3) Single family residential area and locations with Levels 1/2 chargers have a higher start SOC and longer dwell time compared with other locations and DC fast chargers (DCFCs). These findings provide policy insights to identify potential time and locations to incentivize PEVs for grid services, as well as identify critical location categories for further charging infrastructure investment to better reduce range anxiety and promote PEV adoption.

33 ADVANCED PROPULSION SYSTEMS↗

Enriching the Twitter Stream Increasing Data Mining Yield and Quality Using Machine Learning

Social media data streams are important sources of real-time and historical global information for science applications. At the NASA Goddard Earth Sciences Data and Information Services Center (GES DISC), we are exploring the Twitter data stream for its potential in augmenting the validation program of NASA Earth science missions, specifically the Global Precipitation Measurement (GPM) mission. We have implemented a tweet processing infrastructure that outputs classified precipitation tweets. Inputs are "passive" tweets, along with a smaller number of tweets from "active" participants, i.e., those knowingly contributing to our effort. The "active" tweets, presumably of higher quality, enrich the Twitter stream. "Active" sources include data scraped from other social media (e.g., public Facebook posts) and data from existing crowdsourcing programs (e.g., mPING reports). In addition, there is likely relevant precipitation information in images and documents that are the end points of links often included in tweets. Information derived from these "active" sources could then be tweeted into the Twitter stream, thus enriching its quality. The objective of our current work is to mine these tweet­ linked images and documents, using neural networks, to increase the information content and quality related to precipitation. For images, we classified them as either precipitation-related or not. For training and validation, we used images obtained via the Google custom search API. We created two models: (1) by training a simple Convolutional Neural Network and (2) by using transfer learning principles to adapt a pre-trained object recognition model. For documents, both those linked to tweets and the tweet contents, we trained Hierarchical Attention Networks to determine precipitation occurrence, type, and intensity. For training and validation, we used a keyword-filtered tweet data set labelled with ground truth data from Dark Sky (an API to retrieve weather-related labels) and the National Severe Storms Laboratory's Multi­ Radar/Multi-Sensor (MRMS) system. Our results demonstrated the efficacy of our machine learning approaches for enriching the Twitter stream, to derive information potentially useful for validation of earth science satellite data.

Albayrak, Arif↗

Community Data Mining Approach for Surface Complexation Database Development

This paper presents a comprehensive data-to-model workflow, including a findable, accessible, interoperable, reusable (FAIR) community sorption database (newly developed LLNL Surface Complexation/Ion Exchange (L-SCIE) database) along with a data fitting workflow to efficiently optimize surface complexation reaction constants with multiple surface complexation model (SCM) constructs. This workflow serves as a universal framework to mine, compile, and analyze large numbers of published sorption data as well as to estimate reaction constants for parameterizing reactive transport models. Here the framework includes (1) data digitization from published papers, (2) data unification including unit conversions, and (3) data-model integration and reaction constant estimation using geochemical software PHREEQC coupled with the universal parameter estimation code PEST. We demonstrate our approach using an analysis of U(VI) sorption to quartz based on a first L-SCIE implementation, concluding that a multisite SCM construct with carbonate surface species yielded the best fit to community data. Surface complexation reaction constants extracted from this approach captured all available sorption data available in the literature and provided insight into previously published reaction constants and surface complexation model constructs. The L-SCIE sorption database presented herein allows for automating this approach across a wide range of metals and minerals and implementing novel machine learning approaches to reactive transport in the future.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

Data Mining of Groundwater to Identify MAGs with Methane, Propane and Toluene Monooxygenases

Whole genome sequencing datasets, involving more than 600 groundwater samples, from nine countries, were analyzed to identify metagenome assembled genomes (MAGs) containing full operons for propane monooxygenase, soluble methane monooxygease, toluene monooxygenase and particulate ammonia/methane monooxygenase. The enzymes encoded by these genes are a focus of interest because of their ability to degrade common groundwater contaminants. Due to the large amount of data, sequence analyses involved more than 80 individual KBase narratives. The approach followed the KBase tutorial called "Metagenome-Assembled Genome Extraction from a Compost Microbiome Enrichment" The generated MAGs were exported from each individual narrative into separate summary KBase narratives for each monooxygenase. Three KBase narratives were generated for particulate ammonia/methane monooxygenase, due to the large number of MAGs identified.

59 BASIC BIOLOGICAL SCIENCES↗

My vehicle is a data mine

In this talk we explore how analysis of vehicle data provides information of individual vehicle behaviors, information of other vehicles in the flow of traffic, and insights into the behavior of drivers. Over the last two decades traditional passenger vehicles have been transformed from integrated two-port electrical nodes to cyber-physical systems of communicating computational nodes whose individual state and control variables are shared on a standard controller area network (CAN) bus. As driver assistance systems have crept into vehicles as safety features, driver behaviors can be observed through analysis of the data streams on the CAN bus as these new nodes communicate with one another. The properties of these data streams, as well as architectures and approaches to gather the data, are important to consider when drawing conclusions on the relevance of the data in making decisions at varying levels of the information hierarchy. We will demonstrate several technical challenges associated with these data collection processes, as well as preliminary results that demonstrate application relevance of the data to behavior, traffic, and systems domains.

42 ENGINEERING↗

Unsupervised atomic data mining via multi-kernel graph autoencoders for machine learning force fields

Constructing a chemically diverse dataset while avoiding sampling bias is critical to training efficient and generalizable force fields. However, in computational chemistry and materials science, many common dataset generation techniques are prone to oversampling regions of the potential energy surface. Furthermore, these regions can be difficult to identify and isolate from each other or may not align well with human intuition, making it challenging to systematically remove bias in the dataset. While traditional clustering and pruning (down-sampling) approaches can be useful for this, they can often lead to information loss or a failure to properly identify distinct regions of the potential energy surface due to difficulties associated with the high dimensionality of atomic descriptors. In this work, we introduce the Multi-kernel Edge Attention-based Graph Autoencoder (MEAGraph) model, an unsupervised approach for analyzing atomic datasets. MEAGraph combines multiple linear kernel transformations with attention-based message passing to capture geometric sensitivity and enable effective dataset pruning without relying on labels or extensive training. Demonstrated applications on niobium, tantalum, and iron datasets show that MEAGraph efficiently groups similar atomic environments, allowing for the use of basic pruning techniques for removing sampling bias. This approach provides an effective method for representation learning and clustering that can be used for data analysis, outlier detection, and dataset optimization.

Materials science↗

Grids for Dummies: Featuring Earth Science Data Mining Application

This viewgraph presentation discusses the concept and advantages of linking computers together into data grids, an emerging technology for managing information across institutions, and potential users of data grids. The logistics of access to a grid, including the use of the World Wide Web to access grids, and security concerns are also discussed. The potential usefulness of data grids to the earth science community is also discussed, as well as the Global Grid Forum, and other efforts to establish standards for data grids.

Hinke, Thomas H.↗

Grist : grid-based data mining for astronomy

The Grist project is developing a grid-technology based system as a research environment for astronomy with massive and complex datasets. This knowledge extraction system will consist of a library of distributed grid services controlled by a workflow system, compliant with standards emerging from the grid computing, web services, and virtual observatory communities. This new technology is being used to find high redshift quasars, study peculiar variable objects, search for transients in real time, and fit SDSS QSO spectra to measure black hole masses. Grist services are also a component of the 'hyperatlas' project to serve high-resolution multi-wavelength imagery over the Internet. In support of these science and outreach objectives, the Grist framework will provide the enabling fabric to tie together distributed grid services in the areas of data access, federation, mining, subsetting, source extraction, image mosaicking, statistics, and visualization.

grid computing↗

Monitoring Bone Health after Spaceflight: Data Mining to Support an Epidemiological Analysis of Age-related Bone Loss in Astronauts

Through the epidemiological analysis of bone data, HRP is seeking evidence as to whether the prolonged exposure to microgravity of low earth orbit predisposes crewmembers to an earlier onset of osteoporosis. While this collaborative Epidemiological Project may be currently limited by the number of ISS persons providing relevant spaceflight medical data, a positive note is that it compares medical data of astronauts to data of an age-matched (not elderly) population that is followed longitudinally with similar technologies. The inclusion of data from non-ISS and non-NASA crewmembers is also being pursued. The ultimate goal of this study is to provide critical information for NASA to understand the impact of low physical or minimal weight-bearing activity on the aging process as well as to direct its development of countermeasures and rehabilitation programs to influence skeletal recovery. However, in order to optimize these results NASA needs to better define the requirements for long term monitoring and encourage both active and retired astronauts to contribute to a legacy of data that will define human health risks in space.

Baker, K. S,↗

TRMM Data Mining Service at the Goddard Earth Sciences (GES) DISC DAAC Tropical Rainfall Measuring Mission (TRMM)

TRMM has acquired more than four years of data since its launch in November 1997. All TRMM standard products are processed by the TRMM Science Data and Information System (TSDIS) and archived and distributed to general users by the GES DAAC. Table 1 shows the total archive and distribution as of February 28, 2002. The Utilization Ratio (UR), defined as the ratio of the number of distributed files to the number of archived files, of the TRMM standard products has been steadily increasing since 1998 and is currently at 6.98.

Source record↗