Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Topic Models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Natural Language Processing Techniques for Intelligent Knowledge Management of Safety Reports

Safety, failure, and incident reports are common artifacts across various domains, including aviation and wildfire response. These reports are often mandatory to submit, resulting in the culmination of large repositories of text-based documents. Simultaneously, these reports and corresponding repositories are often only manually analyzed and queried by users via out-of-date search engines. As a consequence, we have been developing the Manager for Intelligent Knowledge Access (MIKA) toolkit, which uses natural language processing to improve information access and reuse. In this presentation, we discuss natural language processing techniques for knowledge discovery and apply these methods to a repository of aerial wildfire mishap reports. Two methods are used for knowledge discovery: topic modeling and named-entity recognition. We use topic modeling to identify hazards and perform a trend analysis to produce a data-driven risk matrix. A custom named-entity recognition model, build from fine tuning a pre-trained language model, is used to identify failure modes, failure causes, failure effects, control processes, and recommendations to aid in failure modes and effects analysis (FMEA). Throughout the presentation, we discuss and apply natural language processing techniques to better leverage the vast amount of information contained in report repositories.

Machine learning↗

Improved large perturbation propulsion models for control system design (1988-1989) and large perturbation models of high velocity propulsion systems (1989-1990) and reduced order propulsion models for control system design (1990-1991)

Methods for modeling high speed propulsion systems will be discussed. Included in this category are internal flow propulsion systems without rotating machinery, such as inlets, ramjets, and scramjets. Among the modeling topics discussed are modeling of linear isentropic flow, heat exchange, gasdynamics, lumped parameter systems, and infinite dimensional systems. Furthermore, a generalized overview of modeling high speed propulsion systems is presented in this collection of papers.

Hartley, Tom T.↗

Inverse Reinforcement Learning based Bayesian Goal Inference Method for Early Nuclear Proliferation Detection

Traditional methods for detection of nuclear proliferation indicators are usually applied after nuclear proliferation has already occurred. There is a need to advance these methods to perform early detection of nuclear proliferation indicators. In this project, we formulated an early detection problem as a sequential, decision-making, goal inference problem based on research publications of authors, to determine whether it is possible to infer whether an author will publish on a research activity before it has occurred. To develop and test our approach, we selected a civil nuclear activity for our case study. We constructed a state-action-state transition graph from publications of authors associated with the activity and the co-authors of their publications, using titles, abstracts, and author publication sequences. We then used inverse reinforcement learning to model the goal-directed behavior of authors in trajectories that terminate at selected goal states. Using a Bayesian formulation, we computed the probability that authors would reach each selected state from partially observed trajectories of their state transitions in their research topic space. The state with the highest probability was selected as the most probable goal state. Based on our results, we found that 60% of the times we can infer the correct goal state early; sometimes the inference is either delayed, or multiple states could be inferred as goal states. Overall, our results show that it is possible to perform early detection of research activities of authors in a nuclear technology area. Further research is necessary to establish a more accurate understanding of how topic modeling, topic space grid discretization, and the extent of overlap among trajectories of different goal states, affect the goal inference results. The methods developed in this work may be used to enhance data-driven methods for early detection of nuclear proliferation indicators.

97 MATHEMATICS AND COMPUTING↗

ML for microbiomes

The software provides machine learning analysis and visualization to detect patterns in microbiome data, including topic modeling, probabilistic graphical modeling, conventional machine learning methods, and deep learning. The software is written in python and R, it uses some python and R libraries as well as big open-source libraries like sklearn, networkX, pytorch (python), pgmpy (python), and bnlearn (R). It also has a script to use for MALLET and DTM (open-source packages for topic modeling, written in Java).

Kim, Anastasiia↗

Learning Global Proliferation Expertise Evolution Using AI-Driven Analytics and Public Information

Detecting and anticipating global proliferation expertise and capability evolution from unstructured, noisy, and incomplete public data streams is a highly desired, but extremely challenging task. Here, in this article, we present our pioneering data-driven approach to support the non-proliferation mission to detect and explain the evolution of proliferation expertise and capability development globally from terabytes of publicly available information (PAI), focusing on our knowledge extraction pipeline and descriptive analytics. We first discuss how we fuse nine open-source data streams, including multilingual data, to convert 4 TB of unstructured data to structured knowledge and encode dynamically evolving proliferation expertise representations—content and context graphs. For this, we rely on natural language processing (NLP) and deep learning (DL) models to perform information extraction, topic modeling, and distributed text representation (aka embedding) learning. We then present interactive, usable, and explainable descriptive analytics to refine domain knowledge and present it in a human-understandable form. Finally, we introduce future work avenues that will leverage our dynamic knowledge representations and descriptive analytics to enable predictive and prescriptive inferences to achieve real-time domain understanding and contextual reasoning about global proliferation expertise and capability evolution.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

From BERTopic to SysML: Informing Model-Based Failure Analysis With Natural Language Processing for Complex Aerospace Systems

The development of emerging complex aerospace systems will require new approaches for capturing safety incident scenarios as early as possible in the design phase. However, for novel systems, relevant data available is limited. In this work, we propose a framework informing model-based mission assurance activities with historical incident reports, lessons learned, or other relevant engineering documents using natural language processing. In doing so, we investigate whether there is useful information in data sets that are relevant, if not identical, to the system under design and whether, through rigorous systems engineering practice, this information can be effectively leveraged through model-based failure analysis. In a worked case study, we apply state-of-the-art topic modeling techniques to two data sets, a mission relevant data set and a system relevant data set. The sets of topics are merged and interpreted to form a preliminary list of failure topics that can be used to inform the identification of off-nominal modes in the model-based failure modes and effects analysis development. Once data from the system in operation is available, it can be used to update the topics identified. By extracting information about likely failures from relevant historical data sets and utilizing model-based mission assurance to ensure relevance and rigor, unanticipated failures can be reduced, and projects can more effectively learn from past missions.

Failure Analysis↗

From BERTopic to SysML: Informing Model-Based Failure Analysis With Natural Language Processing for Complex Aerospace Systems

The development of emerging complex aerospace systems will require new approaches for capturing safety incident scenarios as early as possible in the design phase. However, for novel systems, relevant data available is limited. In this work, we propose a framework informing model-based mission assurance activities with historical incident reports, lessons learned, or other relevant engineering documents using natural language processing. In doing so, we investigate whether there is useful information in data sets that are relevant, if not identical, to the system under design and whether, through rigorous systems engineering practice, this information can be effectively leveraged through model-based failure analysis. In a worked case study, we apply state-of-the-art topic modeling techniques to two data sets, a mission relevant data set and a system relevant data set. The sets of topics are merged and interpreted to form a preliminary list of failure topics that can be used to inform the identification of off-nominal modes in the model-based failure modes and effects analysis development. Once data from the system in operation is available, it can be used to update the topics identified. By extracting information about likely failures from relevant historical data sets and utilizing model-based mission assurance to ensure relevance and rigor, unanticipated failures can be reduced, and projects can more effectively learn from past missions.

Failure Analysis↗

The present state and the future direction of eddy viscosity models

Information is given in viewgraph form on the present state and future direction of eddy viscosity models. Topics covered include the eddy viscosity dilemma, two-equation models, equations of motion, free shear flows, incompressible free shear flows, model-predicted boundary layer structure, defect layer analysis, the effects of pressure gradients, viscous sublayer structure, wall functions and viscous damping, viscous damping for kappa-omega, the effects of compressibility, perturbation analysis of the wall layer, an alternative compressibility term, unsteady boundary layers, incompressible separation, backstep results, and compressible separation.

Wilcox, David C.↗

Summary of photovoltaic system performance models

A detailed overview of photovoltaics (PV) performance modeling capabilities developed for analyzing PV system and component design and policy issues is provided. A set of 10 performance models are selected which span a representative range of capabilities from generalized first order calculations to highly specialized electrical network simulations. A set of performance modeling topics and characteristics is defined and used to examine some of the major issues associated with photovoltaic performance modeling. Each of the models is described in the context of these topics and characteristics to assess its purpose, approach, and level of detail. The issues are discussed in terms of the range of model capabilities available and summarized in tabular form for quick reference. The models are grouped into categories to illustrate their purposes and perspectives.

Smith, J. H.↗

A toy terrestrial carbon flow model

A generalized carbon flow model for the major terrestrial ecosystems of the world is reported. The model is a simplification of the Century model and the Forest-Biogeochemical model. Topics covered include plant production, decomposition and nutrient cycling, biomes, the utility of the carbon flow model for predicting carbon dynamics under global change, and possible applications to state-and-transition models and environmentally driven global vegetation models.

Parton, William J.↗

Knowledge Discovery for Early Failure Assessment of Complex Engineered Systems Using Natural Language Processing

Emerging complex engineered systems may have unexpected safety issues due to novel operational environments, increasing autonomy, human-machine interaction, and other factors. To prevent failures in operation or testing that necessitate costly redesign, it is desirable to predict likely failure modes early in the design process. Information about past engineering failures in natural language format presents one possible solution by enabling the retrieval of information that can inform new designs. However, identifying documents containing usable information and extracting the required information can be prohibitively time-consuming when implemented at scale. In this research, an automated natural language processing (NLP) framework is proposed to discover relevant knowledge from documents containing failure-related design information. The framework is applied to NASA’s Lessons Learned Information System (LLIS),which is publicly available. Documents containing usable information are filtered using two different NLP-based models. Next, from the identified usable documents, a failure taxonomy is extracted using a partitioned hierarchical topic modeling approach. Partitions of the document describe different sections of the failure taxonomy – i.e., failure, cause of failure, and recommendations – as indicated by the structure of the original document. The extracted failure taxonomy can be leveraged in early design failure assessment methods. Moreover, the framework can be used to identify documents containing usable failure-related design information from other databases and extract relevant information from these documents.

Documentation and Information Science↗

Knowledge Discovery for Early Failure Assessment of Complex Engineered Systems Using Natural Language Processing

Emerging complex engineered systems may have unexpected safety issues due to novel operational environments, increasing autonomy, human-machine interaction, and other factors. To prevent failures in operation or testing that necessitate costly redesign, it is desirable to predict likely failure modes early in the design process. Information about past engineering failures in natural language format presents one possible solution by enabling the retrieval of information that can inform new designs. However, identifying documents containing usable information and extracting the required information can be prohibitively time-consuming when implemented at scale. In this research, an automated natural language processing (NLP) framework is proposed to discover relevant knowledge from documents containing failure-related design information. The framework is applied to NASA’s Lessons Learned Information System (LLIS),which is publicly available. Documents containing usable information are filtered using two different NLP-based models. Next, from the identified usable documents, a failure taxonomy is extracted using a partitioned hierarchical topic modeling approach. Partitions of the document describe different sections of the failure taxonomy – i.e., failure, cause of failure, and recommendations – as indicated by the structure of the original document. The extracted failure taxonomy can be leveraged in early design failure assessment methods. Moreover, the framework can be used to identify documents containing usable failure-related design information from other databases and extract relevant information from these documents.

Documentation and Information Science↗

Correlated Topics in a Scalable Multidimensional Text Cube: Algorithms and Aviation Safety Case Study

As world-wide air traffic continues to grow even at a modest pace, the overall complexity of the system will increase significantly. This increased complexity can lead to a larger number of fatalities per year even if the extremely low fatality rate that we currently enjoy is maintained. One important source of information about the safety of the aviation system is in Aviation Safety Text Reports which are written by members of the flight crew, air traffic controllers, and other parties involved with the aviation system. These anonymized narrative reports contain fixed-field contextual information about the flight but also contain free-form narratives that describe, in the author s own words, the nature of the safety incident and, in many cases, the contributing factors that led to the safety incident. Several thousand such reports are filed each month, each of which is read and analyzed by highly trained experts. However, it is possible that there are emerging safety issues due to the fact that they may be reported very infrequently and in different contexts with different descriptions. The goal of this research paper is to develop correlated topic models which uncover correlations in the subspaces defined by the intersection of numerous fixed fields and discovered correlated topics. This task requires the discovery of latent topics in the text reports and the creation of a topic cube. Furthermore, because the number of potential cells in the topic cube is very large, we discuss novel methods of pruning the search space in the topic cells, thereby making the analysis feasible. We demonstrate the new algorithms on an analysis of pilot fatigue and its contributing factors, as well as the safety incidents that are correlated with this phenomenon.

Zhao, Bo↗

Investigation of models for large-scale meteorological prediction experiments

The feasibility of extended and long-range weather prediction by means of global atmospheric models was studied. A number of computer experiments were conducted at GISS with the GISS global general circulation model. Topics discussed include atmospheric response to sea-surface temperature anomalies, and monthly mean forecast experiments with the global model.

Spar, J.↗

“Thought I’d Share First” and Other Conspiracy Theory Tweets from the COVID-19 Infodemic: Exploratory Study

Background: The COVID-19 outbreak has left many people isolated within their homes; these people are turning to social media for news and social connection, which leaves them vulnerable to believing and sharing misinformation. Health-related misinformation threatens adherence to public health messaging, and monitoring its spread on social media is critical to understanding the evolution of ideas that have potentially negative public health impacts. Objective: The aim of this study is to use Twitter data to explore methods to characterize and classify four COVID-19 conspiracy theories and to provide context for each of these conspiracy theories through the first 5 months of the pandemic. Methods: We began with a corpus of COVID-19 tweets (approximately 120 million) spanning late January to early May 2020. We first filtered tweets using regular expressions (n=1.8 million) and used random forest classification models to identify tweets related to four conspiracy theories. Our classified data sets were then used in downstream sentiment analysis and dynamic topic modeling to characterize the linguistic features of COVID-19 conspiracy theories as they evolve over time. Results: Analysis using model-labeled data was beneficial for increasing the proportion of data matching misinformation indicators. Random forest classifier metrics varied across the four conspiracy theories considered (F1 scores between 0.347 and 0.857); this performance increased as the given conspiracy theory was more narrowly defined. We showed that misinformation tweets demonstrate more negative sentiment when compared to non-misinformation tweets and that theories evolve over time, incorporating details from unrelated conspiracy theories as well as real-world events. Conclusions: Although we focus here on health-related misinformation, this combination of approaches is not specific to public health and is valuable for characterizing misinformation in general, which is an important first step in creating targeted messaging to counteract its spread. Initial messaging should aim to preempt generalized misinformation before it becomes widespread, while later messaging will

5g↗

The evaluation of a shuttle borne lidar experiment to measure the global distribution of aerosols and their effect on the atmospheric heat budget

A shuttle-borne lidar system is described, which will provide basic data about aerosol distributions for developing climatological models. Topics discussed include: (1) present knowledge of the physical characteristics of desert aerosols and the absorption characteristics of atmospheric gas, (2) radiative heating computations, and (3) general circulation models. The characteristics of a shuttle-borne radar are presented along with some laboratory studies which identify schemes that permit the implementation of a high spectral resolution lidar system.

Shipley, S. T.↗