Engineering PapersSearch

SEARCH · Engineering Papers

Results for “text mining”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Understanding the International Space Station Crew Perspective following Long-Duration Missions through Data Analytics & Visualization of Crew Feedback

The International Space Station (ISS) first became a home and research laboratory for NASA and International Partner crewmembers over 16 years ago. Each ISS mission lasts approximately 6 months and consists of three to six crewmembers. After returning to Earth, most crewmembers participate in an extensive series of 30+ debriefs intended to further understand life onboard ISS and allow crews to reflect on their experiences. Examples of debrief data collected include ISS crew feedback about sleep, dining, payload science, scheduling and time planning, health & safety, and maintenance. The Flight Crew Integration (FCI) Operational Habitability (OpsHab) team, based at Johnson Space Center (JSC), is a small group of Human Factors engineers and one stenographer that has worked collaboratively with the NASA Astronaut office and ISS Program to collect, maintain, disseminate and analyze this data. The database provides an exceptional and unique resource for understanding the "crew perspective" on long duration space missions. Data is formatted and categorized to allow for ease of search, reporting, and ultimately trending, in order to understand lessons learned, recurring issues and efficiencies gained over time. Recently, the FCI OpsHab team began collaborating with the NASA JSC Knowledge Management team to provide analytical analysis and visualization of these over 75,000 crew comments in order to better ascertain the crew's perspective on long duration spaceflight and gain insight on changes over time. In this initial phase of study, a text mining framework was used to cluster similar comments and develop measures of similarity useful for identifying relevant topics affecting crew health or performance, locating similar comments when a particular issue or item of operational interest is identified, and providing search capabilities to identify information pertinent to future spaceflight systems and processes for things like procedure development and training. In addition, the comments were scored for sentiment using a polarity scoring algorithm to identify both positive and negative comments for particular groups and clusters, allowing the team to make analytically informed decisions regarding future hardware and operating procedures. The use of polarity scoring with time series analysis was used to provide insight into how crew health and habitability is changing throughout various spaceflight increments or the station lifecycle as a whole. Finally, a visualization framework was developed to address the needs of the end users to search for and analyze comments by user, category or mission. This paper will discuss how the use of an analytical framework in conjunction with the current human interface, improved the understanding of crew perspective and shortened the time for analysis allowing for more informed decisions and rapid development of improvements. These methods are significantly optimizing the way that this valuable data can be assessed and applied to current and future spaceflight design and development. This collaboration allows the FCI OpsHab team to effectively analyze and share data in a more automated and timely fashion. Trends are no longer derived manually and can be illustrated effectively and accurately with these evolving techniques to an ever growing group of human spaceflight end users.

Bryant, Cody

Mining and beneficiation: A review of possible lunar applications

Successful exploration of Mars and outer space may require base stations strategically located on the Moon. Such bases must develop a certain self-sufficiency, particularly in the critical life support materials, fuel components, and construction materials. Technology is reviewed for the first steps in lunar resource recovery-mining and beneficiation. The topic is covered in three main categories: site selection; mining; and beneficiation. It will also include (in less detail) in-situ processes. The text described mining technology ranging from simple diggings and hauling vehicles (the strawman) to more specialized technology including underground excavation methods. The section of beneficiation emphasizes dry separation techniques and methods of sorting the ore by particle size. In-situ processes, chemical and thermal, are identified to stimulate further thinking by future researchers.

Chamberlain, Peter G.

Agent-Based, Bottom-Up Medium- and Heavy-duty Electric Vehicle Economics, Operation, Charging and Adoption (Research Performance Final Report)

This is the research performance final report for the project entitled: Agent-Based, Bottom-Up Medium- and Heavy-duty Electric Vehicle Economics, Operation, Charging and Adoption This project was able to achieve the DOE’s goals of developing new modeling tools to understand MDHD vehicle operation and adoption. The first modeling tool is a fleet-level techno-economic analysis model capable of estimating energy use and associated environmental and cost impacts for electrified and conventional vehicles of any MDHD vocation, using real-world cost and operations data, including approaches to optimizing schedules for charging and/or vehicle dispatch. The second modeling tool is a system-level, bottom-up, agent-based adoption model capable of generating geographically-resolved estimates of market projections for MDHD vehicles and charging infrastructure. These tools will be developed and published to serve dual purposes as analysis tools for researchers, and decision-support tools for decision makers within the MDHD system.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Assessing the use of UAS-related terms in ASRS using Seed Topic Modeling

Context: The NASA Aviation Safety Reporting System (ASRS) is a voluntary confidential aviation safety reporting system. The ASRS receives reports from pilots, air traffic controllers, flight attendants and other involved in aviation operations. The reports are de-identified and coded by ASRS expert safety analysts. The de-identified reports are then disseminated to the aviation community in a number of ways including entry into an online database. Augmenting the discovery of topics of user interest in this online database would therefore be beneficial to the community it serves. Aim: We propose and execute an experiment to assess the use of seed term topic modeling using the database narratives to identify UAS reports. The use of seed term topic modeling would enable users to identify groups of related narratives associated to a topic of their interest. Method: We use a newly curated field in ASRS reports which identify UAS from non-UAS reports in combination of different set of UAS related terms to assess if seed topic modeling can be used in ASRS.

LDA

Assessing the Use of UAS-Related Terms in ASRS Using Seed Topic Modeling

Context: The NASA Aviation Safety Reporting System (ASRS) is a voluntary confidential system that disseminates reports received from personnel involved in aviation operations after de-identifying them. These reports are used by the community to improve overall aviation system safety. Aim: We propose and execute an experiment to assess the use of seed term topic modeling over the database narratives to identify Unmanned Aircraft System (UAS) reports. The use of seed term topic modeling enables users to identify groups of conceptually similar narratives associated to a topic of their interest. Method: We use a collection of narratives, expert-selected words, and report metadata that separates UAS from non-UAS reports to assess if seed topic modeling can be used to improve ASRS searches. Results: For simpler queries, seed topic search observes a higher recall and lower precision than the existing DBOL (DataBase OnLine) search in operation. However, the best results are obtained when seed topic search is used as a search suggestion system to be executed on the DBOL. Conclusion: Utilizing a combination of both the existing method and the proposed method, users can expand their search vocabulary about subjects of interest while improving the quality of results.

Text Mining

Assessing the Use of UAS-Related Terms in ASRS using Seeds for Topic Modeling

Context: The NASA Aviation Safety Reporting System (ASRS) is a voluntary confidential system that disseminates reports received from personnel involved in aviation operations after de-identifying them. These reports are used by the community to improve overall aviation system safety. Aim: We propose and execute an experiment to assess the use of seed term topic modeling over the database narratives to identify Unmanned Aircraft System (UAS) reports. The use of seed term topic modeling enables users to identify groups of conceptually similar narratives associated to a topic of their interest. Method: We use a collection of narratives, expert-selected words, and report metadata that separates UAS from non-UAS reports to assess if seed topic modeling can be used to improve ASRS searches. Results: For simpler queries, seed topic search observes a higher recall and lower precision than the existing DBOL (DataBase OnLine) search in operation. However, the best results are obtained when seed topic search is used as a search suggestion system to be executed on the DBOL. Conclusion: Utilizing a combination of both the existing method and the proposed method, users can expand their search vocabulary about subjects of interest while improving the quality of results.

LDA

An Approach to Identifying Aspects of Positive Pilot Behavior within the Aviation Safety Reporting System

The National Airspace System (NAS) is constantly evolving as air traffic continues to ramp up to pre-pandemic numbers and projected to grow to unprecedented levels in the coming years. As well as increasing demand to the current system, emerging operations such as Unmanned Autonomous Systems are also expected to add to complexity in the airspace. To address these issues, the industry and government agencies supporting the NAS will need to rely upon additional automation and new technologies to address future operational requirements, while continuing to be a world-leading safe transportation system. As these new technologies are implemented, the system continues to rely on human pilots and controllers in the loop to monitor the system and intervene in situations the automation cannot handle. The goal of proactively addressing safety is of foremost concern to ensure passenger confidence. The industry has implemented various Safety Monitoring Systems to identify safety risks and proactively address them before they result in a serious incident or accident. One such program is the Aviation Safety Reporting System (ASRS). ASRS is a long-established system where pilots and controllers voluntarily and anonymously report safety incidents they experienced and observed during line operations by providing rich text narratives describing the events, the environment, and conditions leading to the safety event of concern. These narratives provide insight and context around events of interest and can be used to identify emerging problems. They can trigger investigations within Flight Operational Quality Assurance or Flight Data Monitoring programs. However, this process typically focuses on the adverse events and the unsafe aspects of the operations surrounding the reported or detected events. This perspective of investigating factors that went wrong around an adverse event is commonly referred to as Safety I. Alternatively, characterizing successful actions that operators perform every day under varying conditions that keep the system within safe operating bounds is a concept referred to as Safety II. The benefit of the Safety II view is that the scope is much larger than that of Safety I since a vast majority of the operations result in successful flights. Many of the successful techniques used to manage operational threats are not documented in standard operating procedures or taught during training. They are typically acquired over time by working with experienced pilots during line operations or in many cases after experiencing a problem for the first time and reacting to it in situ, drawing from years of experience to manage the threat. In an attempt to quantify these positive actions, we are proposing an approach to extracting key behaviors within ASRS reports that can support the Safety II concept. Our analysis assumes that ASRS reports contain some descriptions of corrective actions that operators performed to prevent a situation from leading to an accident. Leveraging recent advances in Natural Language Process modeling, we have developed an approach to extract positive sentiment from reports, embed these positive statements in a vector space where they can be numerically analyzed, and clustering these statements into similar contextual categories. From these contextualized categories we can attempt to summarized and distilled aspects of the positive behavior. The goal is to identify categories of behavior that describe consistent operator techniques that supports the Safety II concept. With this information, airlines may enable learning from these positive actions, or address procedures that need to be changed to avoid having pilots implement a workaround. These insights can provide a lens into what is “going right” in the operations that may otherwise not be known widely within the community. It is envisioned that this approach can be extended to other narrative programs such as Line Operation Safety Audit or Learning Improvement Team reports where similar observed behavior can be analyzed to extract positive actions and inform the overall operations.

NLP

DancePartner: Python Package to Mine Multiomics Relationship Networks from Literature and Databases

A goal of multi-omics experiments is to understand how mechanistic molecular biology is altered between conditions, typically a control group and experimental groups. Oftentimes this involves studying changes in biomolecule relationships (e.g. interactions, metabolic relationships) of several types of biomolecules (e.g. proteins, lipids, metabolites). Though several databases contain relationships between biomolecules, understudied species may have little to no relationship information in databases and thus must be mined from literature. There are several challenges to literature mining, including automated full-text extraction, duplicate biomolecule term collapsing, and implementing complex machine learning tools. To make relationship extraction more accessible to the community, a python package called DancePartner was developed to allow for the extraction of relationships from literature and databases, with functions to map biomolecule synonyms to standardized identifiers and visualize and characterize the resulting multi-omics network. Here, in this study, an example dataset involving Caenorhabditis elegans is presented, where relationships are mined from 1443 publications using DancePartner. These relationships are combined with relationships from KEGG, WikiPathways, UniProt, and LipidMaps, and visualized.

BERT

Informing Plant Asset Reliability and Availability Through AI-Driven Analysis of Operator Logs

The availability and reliability of nuclear power plant (NPP) structures, systems, and components (SSCs) are critical parameters for NPP safety. Tracking these parameters is necessary but costly and labor-intensive, requiring the collection and evaluation of SSC event data such as shutdowns, startups, and failures. To show how these events are needed for the parameters an example is given: one measure of reliability is based on the number of equipment failure events and the number of run hours (i.e., the time from a startup event to a shutdown event). Here, this work investigates using artificial intelligence (AI) to mine NPP operator log entry texts for SSC event data. Four AI approaches were explored for identifying these events, including natural language processing (NLP) methods, generative AI, generative AI combined with NLP, and topic modeling. A key challenge addressed with all four approaches is the brevity of operator log entries. Among these four a neural network–based NLP method was shown to be the most promising for this application, achieving F1 scores of 86.0% for shutdowns, 92.2% for startups, and 80.4% for failures on a subject-matter-expert-curated dataset from NPP operator logs, compared to a baseline of 66.6% for a random classifier. This shows that NLP methods can perform better than generative AI. Additionally, the NLP methods combined with generative AI were shown to perform better than generative AI alone. Generative AI was most successful at providing the background information for the NLP methods to use. This work demonstrates the potential to use AI to automate parameter collection from NPP operator log entries and other records.

97 - MATHEMATICS AND COMPUTING

Natural Language Processing to Inform Agent-Based Modeling: With Application to Modeling Adoption of Medium-Duty Electric Vehicles

Agent-based socio-technical modeling of medium- and heavy-duty (MDHD) electric vehicle (EV) adoption has the potential to provide analysis, prediction, and gui. This paper describes new applications of text analysis developed through machine learning (ML) to build and understand relevant topics and their saliency in the published discourse on adoption of MDHD EVs. This work contributes to the state of the art in topic mining models by defining a new metric of topic ranking (START) that quantifies the importance of predefined topics within the corpus using weighted results for predefined topics from two topic modeling approaches: Latent Dirichlet Allocation (LDA) and BERTopic. The START metric is then demonstrated in practice to model how academia and industry view the EV adoption process based on the respective texts published by these groups. Results show that academic literature places more emphasis on categories of interests such as norms/attitudes and adopter knowledge, while trade journals tend to emphasize long-term cost more than academia. The two bodies of literature agree on the importance of policy and incentives in MDHD EV adoption. Together these results illustrate the potential to use ML-based text analysis to populate the characteristics of agent-based socio-technical models.

Electric vehicle adoption, fleet electrification,

Using a Knowledge Graph to Discover Earth Science Information

Knowledge graphs link key entities within a specific domain to other entities via relationships. Researchers are able to mine these relationships from numerous sources to infer new knowledge. Text extraction from peer-reviewed papers and scientific reports are untapped resources that can be leveraged by knowledge graphs to accelerate scientific discovery.

Freitag, Brian

Data Mining SIAM Presentation

This viewgraph document describes the data mining system developed at NASA Ames. Many NASA programs have large numbers (and types) of problem reports.These free text reports are written by a number of different people, thus the emphasis and wording vary considerably With so much data to sift through, analysts (subject experts) need help identifying any possible safety issues or concerns and help them confirm that they haven't missed important problems. Unsupervised clustering is the initial step to accomplish this; We think we can go much farther, specifically, identify possible recurring anomalies. Recurring anomalies may be indicators of larger systemic problems. The requirement to identify these anomalies has led to the development of Recurring Anomaly Discovery System (ReADS).

Srivastava, Ashok

Searching for 'Unknown Unknowns'

The NASA Engineering and Safety Center (NESC) was established to improve safety through engineering excellence within NASA programs and projects. As part of this goal, methods are being investigated to enable the NESC to become proactive in identifying areas that may be precursors to future problems. The goal is to find unknown indicators of future problems, not to duplicate the program-specific trending efforts. The data that is critical for detecting these indicators exist in a plethora of dissimilar non-conformance and other databases (without a common format or taxonomy). In fact, much of the data is unstructured text. However, one common database is not required if the right standards and electronic tools are employed. Electronic data mining is a particularly promising tool for this effort into unsupervised learning of common factors. This work in progress began with a systematic evaluation of available data mining software packages, based on documented decision techniques using weighted criteria. The four packages, which were perceived to have the most promise for NASA applications, are being benchmarked and evaluated by independent contractors. Preliminary recommendations for "best practices" in data mining and trending are provided. Final results and recommendations should be available in the Fall 2005. This critical first step in identifying "unknown unknowns" before they become problems is applicable to any set of engineering or programmatic data.

Parsons, Vickie S.

Exploration Clinical Decision Support System: Medical Data Architecture

The Exploration Clinical Decision Support (ECDS) System project is intended to enhance the Exploration Medical Capability (ExMC) Element for extended duration, deep-space mission planning in HRP. A major development guideline is the Risk of "Adverse Health Outcomes & Decrements in Performance due to Limitations of In-flight Medical Conditions". ECDS attempts to mitigate that Risk by providing crew-specific health information, actionable insight, crew guidance and advice based on computational algorithmic analysis. The availability of inflight health diagnostic computational methods has been identified as an essential capability for human exploration missions. Inflight electronic health data sources are often heterogeneous, and thus may be isolated or not examined as an aggregate whole. The ECDS System objective provides both a data architecture that collects and manages disparate health data, and an active knowledge system that analyzes health evidence to deliver case-specific advice. A single, cohesive space-ready decision support capability that considers all exploration clinical measurements is not commercially available at present. Hence, this Task is a newly coordinated development effort by which ECDS and its supporting data infrastructure will demonstrate the feasibility of intelligent data mining and predictive modeling as a biomedical diagnostic support mechanism on manned exploration missions. The initial step towards ground and flight demonstrations has been the research and development of both image and clinical text-based computer-aided patient diagnosis. Human anatomical images displaying abnormal/pathological features have been annotated using controlled terminology templates, marked-up, and then stored in compliance with the AIM standard. These images have been filtered and disease characterized based on machine learning of semantic and quantitative feature vectors. The next phase will evaluate disease treatment response via quantitative linear dimension biomarkers that enable image content-based retrieval and criteria assessment. In addition, a data mining engine (DME) is applied to cross-sectional adult surveys for predicting occurrence of renal calculi, ranked by statistical significance of demographics and specific food ingestion. In addition to this precursor space flight algorithm training, the DME will utilize a feature-engineering capability for unstructured clinical text classification health discovery. The ECDS backbone is a proposed multi-tier modular architecture providing data messaging protocols, storage, management and real-time patient data access. Technology demonstrations and success metrics will be finalized in FY16.

Biomedical support

First Measurement of Charged Current Muon Neutrino-Induced Kaon Production on Argon

The Micro Booster Neutrino Experiment (MicroBooNE) is a liquid argon time projection chamber (LArTPC) neutrino detector with an 85-ton active volume, located on-axis to the Booster Neutrino Beamline (BNB) at Fermilab. The detector is exposed to a neutrino flux with average energy of 800 MeV and was designed to study neutrino interactions on argon, particularly in the 1 GeV energy range. Among its main goals is the investigation of neutrino-induced strange particle production at final-state, with a specific interest on charged kaons ($K^{+}$). The understanding of this neutrino interaction is crucial for refining background models in future nucleon decay searches on others LArTPCs experiments such as the Deep Underground Neutrino Experiment (DUNE). Additionally, the techniques developed for the identification of neutrino-induced $K^{+}$ will contribute to improving particle identification methods in LArTPC-based detectors. This thesis is reporting the first-ever measurement of the flux-integrated cross section for charged-current muon neutrino-induced $K^{+}$ production on argon, determined to be $7.93 \pm 3.27(\text{stat.}) \pm 2.92(\text{syst.}) \times 10^{-42}$ cm$^2$/nucleon, based on the MicroBooNE dataset corresponding to $6.88 \times 10^{20}$ Protons on Target (POT).

Rodriguez Rondon, Jairo H. [South Dakota Sch. Mine

Mining Product Reviews for Important Product Features of Refurbished iPhones

Problem: Remanufacturers want to increase consumer interest in refurbished products, which motivates the need to understand which product features are important to buyers of refurbished products such as mobile phones. Research Questions: This study addresses two questions. First, which product features are most important for buyers of refurbished iPhones? Second, how do those preferences differ from the preferences of buyers of new iPhones? Methods: Online reviews of iPhones are obtained and converted into a document–term matrix. Using this text model, three subsets of features are identified using statistical analysis of frequency of mention: most frequent, average, and least frequent. A logistic regression (LR) model is then used to identify which features are most predictive of whether a review is for a new or refurbished phone. Results: Buyers of refurbished phones mention battery health, screen/display, shell condition, and brand significantly more often than other features. Directly contrasting reviews of refurbished versus new phones shows that shell condition, brand, speaker, and charger are found to be the most predictive product features indicated in reviews for refurbished phones. Of those, the shell condition is significantly more predictive than the others. Implications: The results identify product features that remanufacturers of iPhones can emphasize to increase customer demand.

Anisi, Atefeh

Complete Decoding and Reporting of Aviation Routine Weather Reports (METARs)

Aviation Routine Weather Report (METAR) provides surface weather information at and around observation stations, including airport terminals. These weather observations are used by pilots for flight planning and by air traffic service providers for managing departure and arrival flights. The METARs are also an important source of weather data for Air Traffic Management (ATM) analysts and researchers at NASA and elsewhere. These researchers use METAR to correlate severe weather events with local or national air traffic actions that restrict air traffic, as one example. A METAR is made up of multiple groups of coded text, each with a specific standard coding format. These groups of coded text are located in two sections of a report: Body and Remarks. The coded text groups in a U.S. METAR are intended to follow the coding standards set by National Oceanic and Atmospheric Administration (NOAA). However, manual data entry and edits made by a human report observer may result in coded text elements that do not follow the standards, especially in the Remarks section. And contrary to the standards, some significant weather observations are noted only in the Remarks section and not in the Body section of the reports. While human readers can infer the intended meaning of non-standard coding of weather conditions, doing so with a computer program is far more challenging. However such programmatic pre-processing is necessary to enable efficient and faster database query when researchers need to perform any significant historical weather analysis. Therefore, to support such analysis, a computer algorithm was developed to identify groups of coded text anywhere in a report and to perform subsequent decoding in software. The algorithm considers common deviations from the standards and data entry mistakes made by observers. The implemented software code was tested to decode 12 million reports and the decoding process was able to completely interpret 99.93 of the reports. This document presents the deviations from the standards and the decoding algorithm. Storing all decoded data in a database allows users to quickly query a large amount of data and to perform data mining on the data. Users can specify complex query criteria not only on date or airport but also on weather condition. This document also describes the design of a database schema for storing the decoded data, and a Data Warehouse web application that allows users to perform reporting and analysis on the decoded data. Finally, this document presents a case study correlating dust storms reported in METARs from the Phoenix International airport with Ground Stops issued by Air Route Traffic Control Centers (ATCSCC). Blowing widespread dust is one of the weather conditions when dust storm occurs. By querying the database, 294 METARs were found to report blowing widespread dust at the Phoenix airport and 41 of them reported such condition only in the Remarks section of the reports. When METAR is a data source for an ATM research, it is important to include weather conditions not only from the Body section but also from the Remarks section of METARs.

METAR Decoder/Parser