Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data mining”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Analyzing Double Delays at Newark Liberty International Airport (EWR)

When weather or congestion impacts the National Airspace System, multiple different Traffic Management Initiatives can be implemented, sometimes with unintended consequences. One particular perceived inequity that is commonly identified is in the interaction between Ground Delay Programs (GDPs) and time based scheduling of internal departures by the Traffic Management Advisor (TMA) (now operationally superseded by the FAA's the Time-Based Flow Management system). Internal departures under TMA scheduling can take large GDP delays, followed by large TMA scheduling delays, because they cannot easily fit into the arrival flow at the runway. In this paper we examine the causes of these double delays through an analysis of arrival operations at Newark Liberty International Airport (EWR) from June to August 2010. TMA scheduling delays are found to be generally higher than TMA airborne metering delays, regardless of prior GDP delays. Depending on how the double delay is defined, between 42 and 62 of all internal departures in GDP and TMA scheduling experienced double delays in this period. A deep dive into the data reveals that contributors to double delays include upstream flights departing before their Expect Departure Clearance Times (EDCTs); differences in the rates used for setting EDCTs and TMA Scheduled Times of Arrival; differences in the arrival demand expected based on EDCTs and the arrival demand entering TMA; and shorter en route times between takeoff and entry into TMA than assumed in the calculation of flight EDCTs, all of which undermine the sequencing and spacing underlying flight EDCTs. Double delays are also found to coincide with periods in which the virtual runway arrival queue being served by a TMA is large, there are periods of high demand relative to capacity, and there are high airborne metering delays. Data mining techniques are used to confirm that each of these factors contribute to the occurrence of double delay andor high internal departure scheduling delay across three months of data from June to August 2010. Predictors of the occurrence of double delay and high TMA scheduling delay are built using logistic regression, providing prediction accuracies of 69 and 73, respectively.

traffic flow management↗

AMIA KDDM Working Group Collaborative Workshop: Enriching Electronic Health Records with Social Determinants of Health to Improve Outcomes and Health Equity

Prior research has demonstrated that social determinants of health (SDoH) are major drivers of health outcomes and contributors to widespread health inequities. It was estimated that, in the United States, SDoH could be responsible for up to 40% of all preventable deaths, significantly higher than the 10-15% for which better medical care is responsible. Public health interventions that target SDoH are instrumental for improving health outcomes and reducing long-standing health inequities. Currently, most mainstream EHR vendors have implemented SDoH screeners in their EHR systems. However, the utility of the screeners is low, rendering patient-level SDoH still widely unavailable in the structured fields. SDoH are sometimes mentioned in free-text clinical notes (e.g., social context section) where natural language processing (NLP) can be applied to extract relevant information. Contextual-level SDoH can be identified from multiple data sources, many of which are publicly available and spatiotemporally linked to EHR data. As such, there is an opportunity for the KDDM research community to create innovative solutions to draw meaningful insights by creating and using rich data with SDoH to improve health outcomes while reducing disparities. In this workshop organized by AMIA Knowledge Discovery and Data Mining Working Group (AMIA KDDM WG), we will invite world-leading experts from academia, national laboratories, and life science industry with varied backgrounds in biomedical informatics, epidemiology, data science, machine learning, natural language processing, and pediatric cardiology to discuss the best practice of capturing, standardizing, and using SDoH information in various applications aiming at improving outcomes and health equity.

He, Zhe↗

Automating Analysis of Neutron Scattering Time-of-Flight Single Crystal Phonon Data

This article introduces software called Phonon Explorer that implements a data mining workflow for large datasets of the neutron scattering function, S(Q, ω), measured on time-of-flight neutron spectrometers. This systematic approach takes advantage of all useful data contained in the dataset. It includes finding Brillouin zones where specific phonons have the highest scattering intensity, background subtraction, combining statistics in multiple Brillouin zones, and separating closely spaced phonon peaks. Using the software reduces the time needed to determine phonon dispersions, linewidths, and eigenvectors by more than an order of magnitude.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

International Space Station (ISS) Anomalies Trending Study: Appendices - Volume II

The NASA Engineering and Safety Center (NESC) set out to utilize data mining and trending techniques to review the anomaly history of the International Space Station (ISS) and provide tools for discipline experts not involved with the ISS Program to search anomaly data to aid in identification of areas that may warrant further investigation. Additionally, the assessment team aimed to develop an approach and skillset for integrating data sets, with the intent of providing an enriched data set for discipline experts to investigate that is easier to navigate, particularly in light of ISS aging and the plan to extend its life into the late 2020s. This document contains the Appendices to the Volume I report.

Beil, Robert J.↗

International Space Station (ISS) Anomalies Trending Study

The NASA Engineering and Safety Center (NESC) set out to utilize data mining and trending techniques to review the anomaly history of the International Space Station (ISS) and provide tools for discipline experts not involved with the ISS Program to search anomaly data to aid in identification of areas that may warrant further investigation. Additionally, the assessment team aimed to develop an approach and skillset for integrating data sets, with the intent of providing an enriched data set for discipline experts to investigate that is easier to navigate, particularly in light of ISS aging and the plan to extend its life into the late 2020s. This report contains the outcome of the NESC Assessment.

Beil, Robert J.↗

A Taxonomic Classification Approach for Global Spatio-temporal Data

The World Bank, World Health Organization, and other major vendors collectively provide thousands of global time series datasets that focus on issues of the environment, public health, economics, violence, education, and national security. Sorting these data into meaningful information requires the use of data mining techniques to cluster trends into an orderly and manageable number of cases. The World SpatioTemporal Analytics and Mapping (WSTAMP) project database (wstamp.ornl.gov) was developed to spatiotemporally harmonize global vendor data (23,300+ attributes, 200+ locations, 50+ years). Within the WSTAMP analytical environment, Dynamic Time Warping (DTW) has been a highly effective data-driven approach for clustering and mapping these time series into national spatiotemporal behavior maps. Two significant properties have surfaced from this work. First, several recognizable cluster patterns have emerged and persist across a range of locations, attributes, and time frames (e.g., increasing, decreasing, rebounding, peak, oscillating). Secondly, practitioners engaging WSTAMP have noted the explanatory and anticipatory value of these patterns and articulated particular interest in detecting them within the spatiotemporal cube. This need was addressed by shifting DTW-based clustering from an open ended, data-driven implementation to a taxonomic pattern matching approach. This paper presents the method including implementation strategies for visualization and human computer interaction and applies the approach to a sample data set and concludes with next steps.

Stewart, Robert↗

Phase Selection Rules of Multi‐Principal Element Alloys

Abstract Computational prediction of phase stability of multi‐principal element alloys (MPEAs) holds a lot of promise for rapid exploration of the enormous design space and autonomous discovery of superior structural and functional properties. Regardless of many plausible works that rely on phenomenological theory and machine learning, precise prediction is still limited by insufficient data and the lack of interpretability of some machine learning algorithms, e.g., convolutional neural network. In this work, a comprehensive approach is presented, encompassing the development of a complete dataset that contains 72 387 density functional theory calculations, as well as a predictive global phenomenological descriptor. The phase selection descriptor, based on atomic electronegativity and valence electron concentration, significantly outperforms the widely used valence electron concentration, excelling in both accuracy (with an f1 score of 63% compared to 47%) and its ability to predict the HCP phase (0.48 recall compared to 0). The comprehensive data mining on the global design space of 61 425 quaternary MPEAs made from 28 possible metals, together with the phenomenological theory and physical interpretation, will set up a solid computational science foundation for data‐driven exploration of MPEAs.

Chemistry↗

MOFX-DB: An Online Database of Computational Adsorption Data for Nanoporous Materials

Machine learning and data mining coupled with molecular modeling have become powerful tools for materials discovery. Metal-organic frameworks (MOFs) are a rich area for this due to their modular construction and numerous applications. Here, we make data from several previous large-scale studies in MOFs and zeolites from our groups (and new data for N 2 and Ar adsorption in MOFs) easily accessible in one place. The database includes over 3 million simulated adsorption data points for H 2 , CH 4 , CO 2 , Xe, Kr, Ar, and N 2 in over 160 000 MOFs and zeolites, textural properties like pore sizes and surface areas, and the structure file for each material. We include metadata about the Monte Carlo simulations to enable reproducibility. The database is searchable by MOF properties, and the data are stored in a standardized JSON format that that is interoperable with the NIST adsorption database. We also identify several MOFs that meet high performance targets for multiple applications, such as high storage capacity for both hydrogen and methane or high CO 2 capacity plus good Xe/Kr selectivity. Here, by providing this data publicly, we hope to facilitate machine learning studies on these materials, leading to new insights on adsorption in MOFs and zeolites.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Utilization of Unsupervised Anomalies Detector as a Tool for Managing the TDRS Constellation at GSFC

NASA’s Goddard Space Flight Center (GSFC) operates a constellation of ten geosynchronous Tracking and Data Relay Satellites (TDRS). The mission of the TDRS constellation is to provide relay communications from low-earth orbiting spacecraft to the primary ground station at the White Sands Complex in Las Cruces, New Mexico. Major customers include the International Space Station and Hubble Space Telescope. The NASA Space Network project office at GSFC manages the constellation of spacecraft. The constellation is over 30 years old, and a wide range of technologies and manufacturing techniques are represented on-orbit. Since 1983, the TDRS constellation has recorded thousands of gigabytes of telemetry data. Spacecraft telemetry data has changed throughout the three generations of TDRS spacecraft, however each spacecraft has the same basic functions with some generational enhancements. The constellation includes several spacecraft that have significantly outlived the manufacturer's projected lifetime. This has provided NASA with a significant benefit in terms of return on investment, however it places a burden on efficient management of the assets for maximum life without permitting a TDRS spacecraft to become stranded in its geosynchronous orbital slot. Consequently, the highest level of attention is paid to systems whose failure could strand a TDRS spacecraft in orbit. In this paper, we proposed two stages of analyzing spacecraft anomalies using data mining (DM) to enhance on-going predictions of spacecraft life, subsystem performance, and analysis of subsystem anomalies. The first stage conducts the unsupervised anomaly detector to detect potential anomalies in real-time telemetry data. The second stage introduced telemetry weight (TW) to each telemetry parameter to determine which parameter caused the strongest anomaly. We will present case studies of some of these analyses and how the data can impact decisions on the management of the constellation.

Ma, Kenneth Y.↗

NASA Software Cost Estimation Model: An Analogy Based Estimation Model

The cost estimation of software development activities is increasingly critical for large scale integrated projects such as those at DOD and NASA especially as the software systems become larger and more complex. As an example MSL (Mars Scientific Laboratory) developed at the Jet Propulsion Laboratory launched with over 2 million lines of code making it the largest robotic spacecraft ever flown (Based on the size of the software). Software development activities are also notorious for their cost growth, with NASA flight software averaging over 50% cost growth. All across the agency, estimators and analysts are increasingly being tasked to develop reliable cost estimates in support of program planning and execution. While there has been extensive work on improving parametric methods there is very little focus on the use of models based on analogy and clustering algorithms. In this paper we summarize our findings on effort/cost model estimation and model development based on ten years of software effort estimation research using data mining and machine learning methods to develop estimation models based on analogy and clustering. The NASA Software Cost Model performance is evaluated by comparing it to COCOMO II, linear regression, and K-­ nearest neighbor prediction model performance on the same data set.

Hihn, Jairus↗

EVA Wiki - Transforming Knowledge Management for EVA Flight Controllers and Instructors

The EVA Wiki was recently implemented as the primary knowledge database to retain critical knowledge and skills in the EVA Operations group at NASA's Johnson Space Center by ensuring that information is recorded in a common, easy to search repository. Prior to the EVA Wiki, information required for EVA flight controllers and instructors was scattered across different sources, including multiple file share directories, SharePoint, individual computers, and paper archives. Many documents were outdated, and data was often difficult to find and distribute. In 2011, a team recognized that these knowledge management problems could be solved by creating an EVA Wiki using MediaWiki, a free and open-source software developed by the Wikimedia Foundation. The EVA Wiki developed into an EVA-specific Wikipedia on an internal NASA server. While the technical implementation of the wiki had many challenges, one of the biggest hurdles came from a cultural shift. Like many enterprise organizations, the EVA Operations group was accustomed to hierarchical data structures and individually-owned documents. Instead of sorting files into various folders, the wiki searches content. Rather than having a single document owner, the wiki harmonized the efforts of many contributors and established an automated revision controlled system. As the group adapted to the wiki, the usefulness of this single portal for information became apparent. It transformed into a useful data mining tool for EVA flight controllers and instructors, as well as hundreds of others that support EVA. Program managers, engineers, astronauts, flight directors, and flight controllers in differing disciplines now have an easier-to-use, searchable system to find EVA data. This paper presents the benefits the EVA Wiki has brought to NASA's EVA community, as well as the cultural challenges it had to overcome.

Johnston, Stephanie S.↗

EVA Wiki - Transforming Knowledge Management for EVA Flight Controllers and Instructors

The EVA (Extravehicular Activity) Wiki was recently implemented as the primary knowledge database to retain critical knowledge and skills in the EVA Operations group at NASA's Johnson Space Center by ensuring that information is recorded in a common, searchable repository. Prior to the EVA Wiki, information required for EVA flight controllers and instructors was scattered across different sources, including multiple file share directories, SharePoint, individual computers, and paper archives. Many documents were outdated, and data was often difficult to find and distribute. In 2011, a team recognized that these knowledge management problems could be solved by creating an EVA Wiki using MediaWiki, a free and open-source software developed by the Wikimedia Foundation. The EVA Wiki developed into an EVA-specific Wikipedia on an internal NASA server. While the technical implementation of the wiki had many challenges, the one of the biggest hurdles came from a cultural shift. Like many enterprise organizations, the EVA Operations group was accustomed to hierarchical data structures and individually-owned documents. Instead of sorting files into various folders, the wiki searches content. Rather than having a single document owner, the wiki harmonized the efforts of many contributors and established an automated revision control system. As the group adapted to the wiki, the usefulness of this single portal for information became apparent. It transformed into a useful data mining tool for EVA flight controllers and instructors, and also for hundreds of other NASA and contract employees. Program managers, engineers, astronauts, flight directors, and flight controllers in differing disciplines now have an easier-to-use, searchable system to find EVA data. This paper presents the benefits the EVA Wiki has brought to NASA's EVA community, as well as the cultural challenges it had to overcome.

Johnston, Stephanie↗

EVA Wiki - Transforming Knowledge Management for EVA Flight Controllers and Instructors

The EVA Wiki was recently implemented as the primary knowledge database to retain critical knowledge and skills in the EVA Operations group at NASA's Johnson Space Center by ensuring that information is recorded in a common, easy to search repository. Prior to the EVA Wiki, information required for EVA flight controllers and instructors was scattered across different sources, including multiple file share directories, SharePoint, individual computers, and paper archives. Many documents were outdated, and data was often difficult to find and distribute. In 2011, a team recognized that these knowledge management problems could be solved by creating an EVA Wiki using MediaWiki, a free and open-source software developed by the Wikimedia Foundation. The EVA Wiki developed into an EVA-specific Wikipedia on an internal NASA server. While the technical implementation of the wiki had many challenges, one of the biggest hurdles came from a cultural shift. Like many enterprise organizations, the EVA Operations group was accustomed to hierarchical data structures and individually-owned documents. Instead of sorting files into various folders, the wiki searches content. Rather than having a single document owner, the wiki harmonized the efforts of many contributors and established an automated revision controlled system. As the group adapted to the wiki, the usefulness of this single portal for information became apparent. It transformed into a useful data mining tool for EVA flight controllers and instructors, as well as hundreds of others that support the EVA. Program managers, engineers, astronauts, flight directors, and flight controllers in differing disciplines now have an easier-to-use, searchable system to find EVA data. This paper presents the benefits the EVA Wiki has brought to NASA's EVA community, as well as the cultural challenges it had to overcome.

Johnston, Stephanie S.↗

Space Weather Impacts to Conjunction Assessment: A NASA Robotic Orbital Safety Perspective

National Aeronautics and Space Administration (NASA) recognizes the risk of on-orbit collisions from other satellites and debris objects and has instituted a process to identify and react to close approaches. The charter of the NASA Robotic Conjunction Assessment Risk Analysis (CARA) task is to protect NASA robotic (unmanned) assets from threats posed by other space objects. Monitoring for potential collisions requires formulating close-approach predictions a week or more in the future to determine analyze, and respond to orbital conjunction events of interest. These predictions require propagation of the latest state vector and covariance assuming a predicted atmospheric density and ballistic coefficient. Any differences between the predicted drag used for propagation and the actual drag experienced by the space objects can potentially affect the conjunction event. Therefore, the space environment itself, in particular how space weather impacts atmospheric drag, is an essential element to understand in order effectively to assess the risk of conjunction events. The focus of this research is to develop a better understanding of the impact of space weather on conjunction assessment activities: both accurately determining the current risk and assessing how that risk may change under dynamic space weather conditions. We are engaged in a data-- ]mining exercise to corroborate whether or not observed changes in a conjunction event's dynamics appear consistent with space weather changes and are interested in developing a framework to respond appropriately to uncertainty in predicted space weather. In particular, we use historical conjunction event data products to search for dynamical effects on satellite orbits from changing atmospheric drag. Increased drag is expected to lower the satellite specific energy and will result in the satellite's being 'later' than expected, which can affect satellite conjunctions in a number of ways depending on the two satellites' orbits and the geometry of the conjunction. These satellite time offsets can form the basis of a new technique under development to determine whether space weather perturbations, such as coronal mass ejections, are likely to increase, decrease, or have a neutral effect on the collision risk due to a particular close approach.

Ghrist, Richard↗

Robust Informatics Infrastructure Required For ICME: Combining Virtual and Experimental Data

With the increased emphasis on reducing the cost and time to market of new materials, the need for robust automated materials information management system(s) enabling sophisticated data mining tools is increasing, as evidenced by the emphasis on Integrated Computational Materials Engineering (ICME) and the recent establishment of the Materials Genome Initiative (MGI). This need is also fueled by the demands for higher efficiency in material testing; consistency, quality and traceability of data; product design; engineering analysis; as well as control of access to proprietary or sensitive information. Further, the use of increasingly sophisticated nonlinear, anisotropic and or multi-scale models requires both the processing of large volumes of test data and complex materials data necessary to establish processing-microstructure-property-performance relationships. Fortunately, material information management systems have kept pace with the growing user demands and evolved to enable: (i) the capture of both point wise data and full spectra of raw data curves, (ii) data management functions such as access, version, and quality controls;(iii) a wide range of data import, export and analysis capabilities; (iv) data pedigree traceability mechanisms; (v) data searching, reporting and viewing tools; and (vi) access to the information via a wide range of interfaces. This paper discusses key principles for the development of a robust materials information management system to enable the connections at various length scales to be made between experimental data and corresponding multiscale modeling toolsets to enable ICME. In particular, NASA Glenn's efforts towards establishing such a database for capturing constitutive modeling behavior for both monolithic and composites materials

Mutli-scale models↗

Living with a Star Space Environment Testbed

Summary of activities: (1) FYO1 NRA - Model development and data mining. (2) FY03 NRA - Flight investigations. (3) SET carrier development. (4) Study for accommodation of SET carrier to support advanced detectors. (5) Collaboration with other programs: LWS TR&T to maximize synergy between TR&T space environment research and SET space environment effects research. LWS Data System to optimize dissemination of SET data. NASA Electronic Parts and Packaging Program to leverage ground testing of technologies. Defense Threat Reduction Agency to leverage ground testing and common interests in advanced detectors. and Air Force Research Laboratory to leverage flight opportunities. (6) Education and Public Outreach.

Barth, Janet↗

Value-added Data Services at the Goddard Earth Sciences Data and Information Services Center

The NASA Goddard Earth Sciences Data and Information Services Center (GES DISC), in addition to serving the Earth Science community as one of the major Distributed Active Archives Centers (DAACs), provides much more than just data. Among the value-added services available to general users are subsetting data spatially and/or by parameter, online analysis (to avoid downloading unnecessarily all the data), and assistance in obtaining data from other centers. Services available to data producers and high-volume users include consulting on building new products with standard formats and metadata and construction of data management systems. A particularly useful service is data processing at the DISC (i.e., close to the input data) with the users algorithm. This can take a number of different forms: as a configuration-managed algorithm within the main processing stream; as a stand-alone program next to the on-line data storage; as build-it-yourself code within the Near-Archive Data Mining (NADM) system; or as an on-the-fly analysis with simple algorithms embedded into the web-based tools. Partnerships between the GES DISC and scientists, both producers and users, allow the scientists to concentrate on science, while the GES DISC handles the data management, e.g., formats, integration, and data processing. The existing data management infrastructure at the GES DISC supports a wide spectrum of options: from simple data support to sophisticated on-line analysis tools, producing economies of scale and rapid time-to-deploy. At the same time, such partnerships allow the GES DISC to serve the user community more efficiently and to better prioritize on-line holdings. Several examples of successful partnerships are described in the presentation.

Leptoukh, Gregory G.↗

An Ensemble Approach to Building Mercer Kernels with Prior Information

This paper presents a new methodology for automatic knowledge driven data mining based on the theory of Mercer Kernels, which are highly nonlinear symmetric positive definite mappings from the original image space to a very high, possibly dimensional feature space. we describe a new method called Mixture Density Mercer Kernels to learn kernel function directly from data, rather than using pre-defined kernels. These data adaptive kernels can encode prior knowledge in the kernel using a Bayesian formulation, thus allowing for physical information to be encoded in the model. Specifically, we demonstrate the use of the algorithm in situations with extremely small samples of data. We compare the results with existing algorithms on data from the Sloan Digital Sky Survey (SDSS) and demonstrate the method's superior performance against standard methods. The code for these experiments has been generated with the AUTOBAYES tool, which automatically generates efficient and documented C/C++ code from abstract statistical model specifications. The core of the system is a schema library which contains templates for learning and knowledge discovery algorithms like different versions of EM, or numeric optimization methods like conjugate gradient methods. The template instantiation is supported by symbolic-algebraic computations, which allows AUTOBAYES to find closed-form solutions and, where possible, to integrate them into the code.

Srivastava, Ashok N.↗