Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Latent Dirichlet Allocation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Resolving Mixtures of Soot Characterized by SP-AMS Spectra Using a Latent Dirichlet Allocation Model

Soot produced by detonation or combustion events exhibits different chemical properties depending on the fuel, device construction, and environmental conditions in which the event occurs. These properties can be useful for defining relevant signatures for probabilistically identifying the different types of events that occurred, based on the soot that is produced from these events. However, it is rare to observe samples of soot from a detonation or combustion that are not contaminated by outside particles. In this paper, we present a method for resolving mixtures of soot to determine the contributions of sources that may be present in samples of recovered soot. We use Latent Dirichlet Allocation to describe the generative process for a sample of recovered soot, and use Variational Bayesian Inference to learn about the parameters associated with the generative model. We demonstrate the utility of this method by considering real samples of mixtures of soot under various frameworks to show that the model is able to identify the different components present in a sample of soot as well as their mixing proportions.

54 ENVIRONMENTAL SCIENCES↗

The Latent Dirichlet Allocation model applied to airborne LiDAR data: A case study on mapping forest degradation associated with fragmentation and fire in the Amazon region

1. LiDAR data are being increasingly used to provide a detailed characterization of the vertical profile of forests. This characterization enables the generation of new insights on the influence of environmental drivers and anthropogenic disturbances on forest structure as well as on how forest structure influences important ecosystem functions and services. Unfortunately, extracting information from LiDAR data in a way that enables the spatial visualization of forest structure, as well as its temporal changes, is challenging due to the high dimensionality of these data. 2. In this study, we show how the Latent Dirichlet Allocation model applied to LiDAR data (LidarLDA) can be used to identify forest structural types and how the relative abundance of these forest types changes throughout the landscape. The code to fit this model is made available through the open-source r package LidarLDA in github. We illustrate the use of LidarLDA both with simulated data and data from a large-scale fire experiment in the Brazilian Amazon region. 3. Using simulated data, we demonstrate that LidarLDA accurately identifies the number of forest types as well as their spatial distribution and absorptance probabilities. For the empirical data, we found that LidarLDA detects both landscape-level patterns in forest structure as well as the strong interacting effect of fire and forest fragmentation on forest structure based on the experimental fire plots. More specifically, LidarLDA reveals that proximity to forest edge exacerbates the impact of fires, and that burned forests remain structurally different from unburned areas for at least 7 years, even when burned only once. Importantly, LidarLDA generates insights on the 3D structure of forest that cannot be obtained using more standard approaches that just focus on top-of-the-canopy information (e.g. canopy height models based on LiDAR data). 4. By enabling the mapping of forest structure and its temporal changes, we believe that LidarLDA will be of broad utility to the ecological research community.

54 ENVIRONMENTAL SCIENCES↗

Latent Dirichlet Allocation modeling of environmental microbiomes

Interactions between stressed organisms and their microbiome environments may provide new routes for understanding and controlling biological systems. However, microbiomes are a form of high-dimensional data, with thousands of taxa present in any given sample, which makes untangling the interaction between an organism and its microbial environment a challenge. Here we apply Latent Dirichlet Allocation (LDA), a technique for language modeling, which decomposes the microbial communities into a set of topics (non-mutually-exclusive sub-communities) that compactly represent the distribution of full communities. LDA provides a lens into the microbiome at broad and fine-grained taxonomic levels, which we show on two datasets. In the first dataset, from the literature, we show how LDA topics succinctly recapitulate many results from a previous study on diseased coral species. We then apply LDA to a new dataset of maize soil microbiomes under drought, and find a large number of significant associations between the microbiome topics and plant traits as well as associations between the microbiome and the experimental factors, e.g. watering level. This yields new information on the plant-microbial interactions in maize and shows that LDA technique is useful for studying the coupling between microbiomes and stressed organisms.

59 BASIC BIOLOGICAL SCIENCES↗

Natural Language Processing to Inform Agent-Based Modeling: With Application to Modeling Adoption of Medium-Duty Electric Vehicles

Agent-based socio-technical modeling of medium- and heavy-duty (MDHD) electric vehicle (EV) adoption has the potential to provide analysis, prediction, and gui. This paper describes new applications of text analysis developed through machine learning (ML) to build and understand relevant topics and their saliency in the published discourse on adoption of MDHD EVs. This work contributes to the state of the art in topic mining models by defining a new metric of topic ranking (START) that quantifies the importance of predefined topics within the corpus using weighted results for predefined topics from two topic modeling approaches: Latent Dirichlet Allocation (LDA) and BERTopic. The START metric is then demonstrated in practice to model how academia and industry view the EV adoption process based on the respective texts published by these groups. Results show that academic literature places more emphasis on categories of interests such as norms/attitudes and adopter knowledge, while trade journals tend to emphasize long-term cost more than academia. The two bodies of literature agree on the importance of policy and incentives in MDHD EV adoption. Together these results illustrate the potential to use ML-based text analysis to populate the characteristics of agent-based socio-technical models.

Electric vehicle adoption, fleet electrification, ↗

Topic Modeling Tool for PeTaL (Periodic Table of Life)

A topic modeling tool is constructed for the purpose of providing insights from biology to the engineer within the framework of PeTaL (Periodic Table of Life). The machine learning text mining tools–latent Dirichlet allocation (LDA) and nonnegative matrix factorization (NMF) with Kullback-Leibler (KL) divergence—are used to provide topic clusters to the user. Topic clusters are the underlying themes of a paper. For the text modeling problem, NMF-KL is the equivalent of probabilistic latent semantic analysis. Both LDA and NMF-KL are top-performing modeling tools. These tools are used to identify biological specimens relevant to the user. Various organisms solve a particular survival problem in nature differently. The topic clusters allow people without domain expertise to find these cross-topic themes in the body of documents and then branch out and examine papers whose target organisms solve the engineer’s problem. Abstracts from the Journal of Experimental Biology were used as input for the clustering tool in addition to a curated set of articles for validation. The tool is able to accept alternate input sources.

Machine learning↗

AI-powered topic modeling: comparing LDA and BERTopic in analyzing opioid-related cardiovascular risks in women

Topic modeling is a crucial technique in natural language processing (NLP), enabling the extraction of latent themes from large text corpora. Traditional topic modeling, such as Latent Dirichlet Allocation (LDA), faces limitations in capturing the semantic relationships in the text document although it has been widely applied in text mining. BERTopic, created in 2022, leveraged advances in deep learning and can capture the contextual relationships between words. In this work, we integrated Artificial Intelligence (AI) modules to LDA and BERTopic and provided a comprehensive comparison on the analysis of prescription opioid-related cardiovascular risks in women. Opioid use can increase the risk of cardiovascular problems in women such as arrhythmia, hypotension etc. 1,837 abstracts were retrieved and downloaded from PubMed as of April 2024 using three Medical Subject Headings (MeSH) words: “opioid,” “cardiovascular,” and “women.” Machine Learning of Language Toolkit (MALLET) was employed for the implementation of LDA. BioBERT was used for document embedding in BERTopic. Eighteen was selected as the optimal topic number for MALLET and 23 for BERTopic. ChatGPT-4-Turbo was integrated to interpret and compare the results. The short descriptions created by ChatGPT for each topic from LDA and BERTopic were highly correlated, and the performance accuracies of LDA and BERTopic were similar as determined by expert manual reviews of the abstracts grouped by their predominant topics. The results of the t-SNE (t-distributed Stochastic Neighbor Embedding) plots showed that the clusters created from BERTopic were more compact and well-separated, representing improved coherence and distinctiveness between the topics. Our findings indicated that AI algorithms could augment both traditional and contemporary topic modeling techniques. In addition, BERTopic has the connection port for ChatGPT-4-Turbo or other large language models in its algorithm for automatic interpretation, while with LDA interpretation must be manually, and needs special procedures for data pre-processing and stop words exclusion. Therefore, while LDA remains valuable for large-scale text analysis with resource constraints, AI-assisted BERTopic offers significant advantages in providing the enhanced interpretability and the improved semantic coherence for extracting valuable insights from textual data.

Research & Experimental Medicine↗

Topic Modeling of NASA Space System Problem Reports: Research in Practice

Problem reports at NASA are similar to bug reports: they capture defects found during test, post-launch operational anomalies, and document the investigation and corrective action of the issue. These artifacts are a rich source of lessons learned for NASA, but are expensive to analyze since problem reports are comprised primarily of natural language text. We apply topic modeling to a corpus of NASA problem reports to extract trends in testing and operational failures. We collected 16,669 problem reports from six NASA space flight missions and applied Latent Dirichlet Allocation topic modeling to the document corpus. We analyze the most popular topics within and across missions, and how popular topics changed over the lifetime of a mission. We find that hardware material and flight software issues are common during the integration and testing phase, while ground station software and equipment issues are more common during the operations phase. We identify a number of challenges in topic modeling for trend analysis: 1) that the process of selecting the topic modeling parameters lacks definitive guidance, 2) defining semantically-meaningful topic labels requires nontrivial effort and domain expertise, 3) topic models derived from the combined corpus of the six missions were biased toward the larger missions, and 4) topics must be semantically distinct as well as cohesive to be useful. Nonetheless,topic modeling can identify problem themes within missions and across mission lifetimes, providing useful feedback to engineers and project managers.

Data Mining↗

Transforming Science Prioritization Processes Using Artificial Intelligence

Artificial Intelligence (AI) and Machine Learning (ML) have potential to augment significantly the current labor-intensive processes of science prioritization, specifically by the National Academies’ Decadal Survey on behalf of NASA and NSF. Here we summarize what we believe to be the first exploratory demonstration-of-concept results from an application of AI/ML to Survey science prioritization. Specifically, we applied Latent Dirichlet Allocation (LDA) and Natural Language Processing (NLP) to reveal trends in published astrophysics research that may indicate science priorities and which could be applied to strategic planning. For the purpose of the work that we summarize here, AI/ML is able to analyze – that is, to “understand,” in a manner of speaking – a vast amount of text to reveal complex relationships among research topics, including the growth or decline of science community activities in those topics over time. We trained ourselves and AI/ML algorithms by using ~400,000 abstracts in the period 1998 to 2010 to “forecast” the Academies’ Astro2010 recommendations and compare with the solicited white papers. Comparing our results with actual Astro2010 recommendations allowed us to identify candidate metrics that better predicted the actual results of the Survey. We found, for example, that Compound Annual Growth Rate (CAGR) of papers published in a topic area is a good proxy measure for importance of this topic area of research. With this training complete, we identified candidate astrophysics astrophysics science priorities for the 2021+ period using the research during 2007 - 2019 . We conclude that appropriate application of AI can potentially significantly reduce the current workload of the Decadal Survey processes and reveal otherwise unrecognized characteristics in the body of astronomical research. We emphasize throughout the exploratory nature of our work, encouraging colleagues to pursue promising results further. Our most critical governing assumption was that increased (or decreased) research activity can be used to identify scientific or technology topic areas worthy of increased (or decreased) future emphasis. We discuss advantages, limitations, and recognize the “black box” nature of our technique. We note ethics issues associated, for example, with using AI/ML to reveal “hidden” meanings and biases in published work. Furthermore, inevitable improvements in AI may soon enable widespread and welcome identification of and advocacy for science and technology priorities by disparate and diverse groups and organizations. Consequently, we continue to urge a near-term, in-depth evaluation of appropriate applications of AI, including implications and consequences, as well as support for multiple follow-on assessments, of which ours is only a beginning.

Artificial Intelligence↗

Exploring the Landscape of Earth and Space Science Informatics using Latent Topic Modeling

AGU Earth and Space Science Informatics (ESSI) is at the forefront of data management, analysis, large scale experimentation, and infrastructure development pertaining to Earth and Space Science interests. The key topics of interest within ESSI are also evolving and diversifying over time. We aim to observe and quantify the various topics covered in ESSI, analyze their trends over time, and identify the contributors’ affiliations to gain an understanding of the landscape of ESSI and the direction of the research and management. The data for this work are abstracts submitted to AGU’s ESSI Fall meeting; They serve as a proxy for key research and development areas within ESSI. We use an unsupervised topic modeling technique called Latent Dirichlet Allocation to observe the underlying topics covered in ESSI and their trends over time. With this presentation, we showcase our results from the analysis and insights gained.

Muthukumaran Ramasubramanian↗

Determining Research Priorities for Astronomy Using Machine Learning

We summarize the first exploratory investigation into whether Machine Learning (ML) techniques can augment science strategic planning. We find that an approach based on Latent Dirichlet Allocation (LDA) using abstracts drawn from high-impact astronomy journals may provide a leading indicator of future interest in a research topic. We show two topic metrics that correlate well with the high-priority research areas identified by the 2010 National Academies’ Astronomy and Astrophysics Decadal Survey. One metric is based on a sum of the fractional contribution to each topic by all scientific papers (“counts”) while the other is the Compound Annual Growth Rate (CAGR) of counts. These same metrics also show the same degree of correlation with the whitepapers submitted to the same Decadal Survey. Our results suggest that the Decadal Survey may under-emphasize fast growing research. A preliminary version of our work was presented by Thronson et al. (2021).

Astronomy↗

A Geometry-Driven Longitudal Topic Model

A simple and scalable framework for longitudinal analysis of Twitter data is developed that combines latent topic models with computational geometric methods. Dimensionality reduction tools from computational geometry are applied to learn the intrinsic manifold on which the latent, temporal topics reside. Then shortest path distances on the manifold are used to link together these topics. The proposed framework permits visualization of the low-dimensional embedding which provides clear interpretation of the complex, high-dimensional trajectories that may exist among latent topics. Practical application of the proposed framework is demonstrated through its ability to capture and effectively visualize natural progression of latent COVID-19 related topics learned from Twitter data. Interpretability of the trajectories is achieved by comparing to real-world events. In addition, the framework permits study of spatial variation in Twitter behavior for learned topics. The analysis demonstrates that the proposed framework is able to capture granular-level impact of COVID-19 on public discussions. We end by arguing that Twitter data, when analyzed within the proposed framework, can serve as a valuable supplementary data stream for COVID-related studies.

97 MATHEMATICS AND COMPUTING↗

Agent-Based, Bottom-Up Medium- and Heavy-duty Electric Vehicle Economics, Operation, Charging and Adoption (Research Performance Final Report)

This is the research performance final report for the project entitled: Agent-Based, Bottom-Up Medium- and Heavy-duty Electric Vehicle Economics, Operation, Charging and Adoption This project was able to achieve the DOE’s goals of developing new modeling tools to understand MDHD vehicle operation and adoption. The first modeling tool is a fleet-level techno-economic analysis model capable of estimating energy use and associated environmental and cost impacts for electrified and conventional vehicles of any MDHD vocation, using real-world cost and operations data, including approaches to optimizing schedules for charging and/or vehicle dispatch. The second modeling tool is a system-level, bottom-up, agent-based adoption model capable of generating geographically-resolved estimates of market projections for MDHD vehicles and charging infrastructure. These tools will be developed and published to serve dual purposes as analysis tools for researchers, and decision-support tools for decision makers within the MDHD system.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗