Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “natural language processing (NLP)”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Vision Transformers Explained Series

Since their introduction in 2017 with Attention is All You Need¹, transformers have established themselves as the state of the art for natural language processing (NLP). In 2021, An Image is Worth 16x16 Words successfully adapted transformers for computer vision tasks. Since then, numerous transformer-based architectures have been proposed for computer vision. This article walks through the Vision Transformer (ViT) as laid out in An Image is Worth 16x16 Words.

97 MATHEMATICS AND COMPUTING↗

Towards Unlocking Insights from Logbooks Using AI

Electronic logbooks contain valuable information about activities and events concerning their associated particle accelerator facilities. However, the highly technical nature of logbook entries can hinder their usability and automation. As natural language processing (NLP) continues advancing, it offers opportunities to address various challenges that logbooks present. This work explores jointly testing a tailored Retrieval Augmented Generation (RAG) model for enhancing the usability of particle accelerator logbooks at institutes like DESY, BESSY, Fermilab, BNL, SLAC, LBNL and CERN. The RAG model uses a corpus built on logbook contributions and aims to unlock insights from these logbooks by leveraging retrieval over facility datasets, including discussion about potential multimodal sources. Our goals are to increase the FAIR-ness (findability, accessibility, interoperability, and reusability) of logbooks by exploiting their information content to streamline everyday use, to enable macro-analysis for root cause analysis, and to facilitate problem-solving automation.

43 PARTICLE ACCELERATORS↗

Adaptive Discovery and Mixed-Variable Optimization of Next Generation Synthesizable Microelectronic Materials

Design of new microelectronic materials is characterized by several challenges such as high-dimensionality of the atomic structure-composition variable space, formidable cost of directly using high-fidelity simulations for design optimization, dispersity in literature-reported similar materials and synthesis methods, complex physical mechanisms, and mixed qualitative and quantitative design variables that lead to a disjointed design space. Even though machine learning (ML) techniques have been employed to expedite materials innovation, existing methods treat ML and design optimization as two separate processes, failing to resolve the fundamental challenges associated with high dimensionality and mixed-variable complexity. We have developed a ML enhanced mixed-variable material design optimization framework to efficiently extract useful information from existing data in literature and physics-based simulations to guide the autonomous search for optimal materials. Our proposed framework is composed of four computational modules: (1) a natural language processing (NLP) based virtual screening module, (2) classification based concept exploration module, (3) a density functional theory (DFT)-based high-fidelity evaluation model, and (4) a novel latent-variable Gaussian process (LVGP) ML model for mixed-variable problems with uncertainty quantification, which seamlessly integrates with Bayesian Optimization (BO) and achieves superb efficiency through embedded physics-based dimension reduction. Our approach is demonstrated and validated using the testbed of functional materials exhibiting metal-insulation transitions (MITs), with the targeted reversible resistivity changes (∼10^5) near room temperature. At the end of the 30-month project, we have developed a series of new ML techniques using NLP, conditional variational autoencoders, active learning, latent-variable Gaussian processes, integrated with Bayesian optimization. Our project has resulted in new predicted MITs compounds and improved understanding of MITs microscopic mechanisms, which in turn will revolutionize microelectronics science to provide energy-saving solutions. Our research has improved both creativity and efficiency in transforming rare-event discoveries of new functional materials to persistent innovations. In addition to open-sourcing the online MIT database and the classification model, the LVGP open source code has been downloaded more than 15,000 times within two years. More than 40 MIT compounds have been identified and many have been pursued experimentally via collaborators. The research results are published in close to 20 collaborative papers in high-impact journals, such as Chem. Mater., Appl. Phys. Rev., Sci. Rep., among others of design space.

36 MATERIALS SCIENCE↗

Knowledge Graph Entity Linking using Graph Embeddings

Details the use of a custom embedding model on knowledge graphs to aid in downstream natural language processing (NLP) models for Derivative Classification Assist. Motivations, algorithms, and results were discussed.

Mahesh, Aarav [Sandia National Laboratories (SNL-N↗

US Hydropower & Environmental Mitigations: A 1998-2023 inventory of mitigation measures inside licensing documents

Federal Energy Regulatory Commission (FERC) license documents mandate environmental mitigation requirements to reduce potential environmental damage caused by non-federal hydropower facilities. This StoryMap showcases a dataset that leveraged Natural Language Processing (NLP) to inventory environmental mitigations from 465 FERC licenses that were issued from 1998-2023 (Ruggles et al, 2025). Users can explore trends in mitigation requirements over time and space with interactive maps and are presented with a demonstration use case.

Ruggles, Thomas A. [Oak Ridge National Laboratory ↗

NukeLM: Pre-Trained and Fine-Tuned Language Models for the Nuclear and Energy Domains

Natural language processing (NLP) tasks (text classification, named entity recognition, etc.) have seen amazing improvements over the last few years. This is due to models such as BERT that achieve deep knowledge transfer by using a large pre-trained model, then fine-tuning the model on specific tasks. The BERT architecture has shown even better performance on domain-specific tasks when the model is pre-trained using domain-relevant texts. Here, inspired by these recent advancements, we have developed NukeLM, a nuclear-domain BERT model pre-trained on 1.5 million abstracts from the DOE Office of Scientific and Technical Information (OSTI) database. This NukeLM model is then fine-tuned for the classification of research articles into either binary classes (related to the nuclear fuel cycle (NFC) or not) or multiple categories related to the subject of the article. We show that continued pre-training of a BERT-style architecture prior to fine-tuning results in greater performance in both article classification tasks. This information is critical for properly triaging manuscripts, a necessary task for better understanding citation networks that publish in the nuclear space and uncovering new areas of research in the nuclear (or nuclear relevant) domain.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

AI-powered topic modeling: comparing LDA and BERTopic in analyzing opioid-related cardiovascular risks in women

Topic modeling is a crucial technique in natural language processing (NLP), enabling the extraction of latent themes from large text corpora. Traditional topic modeling, such as Latent Dirichlet Allocation (LDA), faces limitations in capturing the semantic relationships in the text document although it has been widely applied in text mining. BERTopic, created in 2022, leveraged advances in deep learning and can capture the contextual relationships between words. In this work, we integrated Artificial Intelligence (AI) modules to LDA and BERTopic and provided a comprehensive comparison on the analysis of prescription opioid-related cardiovascular risks in women. Opioid use can increase the risk of cardiovascular problems in women such as arrhythmia, hypotension etc. 1,837 abstracts were retrieved and downloaded from PubMed as of April 2024 using three Medical Subject Headings (MeSH) words: “opioid,” “cardiovascular,” and “women.” Machine Learning of Language Toolkit (MALLET) was employed for the implementation of LDA. BioBERT was used for document embedding in BERTopic. Eighteen was selected as the optimal topic number for MALLET and 23 for BERTopic. ChatGPT-4-Turbo was integrated to interpret and compare the results. The short descriptions created by ChatGPT for each topic from LDA and BERTopic were highly correlated, and the performance accuracies of LDA and BERTopic were similar as determined by expert manual reviews of the abstracts grouped by their predominant topics. The results of the t-SNE (t-distributed Stochastic Neighbor Embedding) plots showed that the clusters created from BERTopic were more compact and well-separated, representing improved coherence and distinctiveness between the topics. Our findings indicated that AI algorithms could augment both traditional and contemporary topic modeling techniques. In addition, BERTopic has the connection port for ChatGPT-4-Turbo or other large language models in its algorithm for automatic interpretation, while with LDA interpretation must be manually, and needs special procedures for data pre-processing and stop words exclusion. Therefore, while LDA remains valuable for large-scale text analysis with resource constraints, AI-assisted BERTopic offers significant advantages in providing the enhanced interpretability and the improved semantic coherence for extracting valuable insights from textual data.

Research & Experimental Medicine↗

AI-powered municipal solid waste management: a comprehensive review from generation to utilization

The accumulation of municipal solid waste (MSW) continues to rise due to burgeoning population, rapid global urbanization and economic growth, intensifying ecological concerns associated with landfills and greenhouse gas (GHG) emissions. Over the past 2 decades, global waste generation has surged by 50%, with one-third remaining uncollected and about 70% sent to landfills. This review examines the critical role of integrating emerging technologies, such as advanced sensors and artificial intelligence (AI), into end-to-end MSW management to alleviate landfill burdens. The suitability of various AI tools for different stages of MSW management is assessed, alongside the deployment of advanced sensors including hyperspectral cameras, computer vision systems, and internet of things (IoT) devices for material identification. Applications of genetic algorithms and reinforcement learning for optimizing collection routes, reducing costs, and lowering emissions are highlighted. Life cycle assessment (LCA) across all stages of MSW management is also reviewed, along with future trends in leveraging generative AI, natural language processing (NLP), and agent-based AI systems to analyze waste generation patterns and public sentiment. Efficient collection and handling can be enhanced through route optimization with geographic information systems and real-time bin-level monitoring. Furthermore, sensor-embedded, real-time object detection systems paired with robotics enable material characterization and automated sorting, thereby lowering costs and diverting waste from landfills into value-added products for diverse industrial sectors including packaging, chemicals, textiles, metals and glass, transportation, and electronics industries. Without intervention, global waste is projected to reach 4.54 billion tons by 2050, contributing direct economic costs of $\$$400 billion and roughly 2.38 billion tons of CO 2 -equivalent emissions annually. This review demonstrates how AI-driven, end-to-end solutions for MSW management can mitigate economic and environmental challenges, while directly supporting the United Nations Sustainable Development (UNDP) goals related to innovation and infrastructure (SDG 9), sustainable cities (SDG 11), responsible consumption and production (SDG 12), and climate action (SDG 13).

09 BIOMASS FUELS↗

Reusing Data and Metadata to Create New Metadata Through Machine-Learning & Other Programmatic Methods

Recent improvements in natural language processing (NLP) enable metadata to be created programmatically from reused original metadata or even the dataset itself. Transfer-learning applied to NLP has greatly improved performance and reduced training data requirements. In this talk, we’ll compare machine-generated metadata to human-generated metadata and discuss characteristics of metadata and data archives that affect suitability for machine-learning reuse of metadata. Where as human-generated metadata is often populated once, populated from the perspective of data supplier, populated by many individuals with different words for the same thing, and limited in length, machine-generated metadata can be updated any number of times, generated from the perspective of any user, constrained to a standardized set of terms that can be evolved over time, and be any length required. Machine-learning generated metadata offers benefits but also additional needs in terms of version control, process transparency, human-computer interaction, and IT requirements. As a successful example, we’ll discuss how a dataset of abstracts and associated human-tagged keywords from a standardized list of several thousand keywords were used to create a machine-learning model that predicted keyword metadata for open-source code projects on code.nasa.gov. We’ll also discuss a less successful example from data.nasa.gov to show how data archive architecture and characteristics of initial metadata can be strong controls on how easy it is to leverage programmatic methods to reuse metadata to create additional metadata.

Gosses, Justin↗

A Hybrid Approach to Labeling Datasets in Earth Science Publications

NASA Data Centers provide the public with thousands of datasets that result in published papers, reports, and conference proceedings. Collecting accurate metrics on usage of these datasets is key to connecting different areas of knowledge and evaluating the datasets’ impact. While most of the datasets have Digital Object Identifiers (DOIs) assigned, most publications do not cite them hampering the automated search of these publications. Instead, articles mention attributes like organization, instrument, mission, variable, or a publication describing the dataset. Often only domain experts can deduce the dataset that was used in the publication text. The lack of a citation slows the spread of information and reduces the research’s impact. With thousands of papers produced each year, an automated means of labeling datasets is critical. This paper explores a hybrid approach of heuristics and a Natural Language Processing (NLP) Named Entity Recognition (NER) model to find and label the datasets used within Earth Science papers. Heuristics are used to produce the labelled sentences and any potential dataset candidates that can be derived from a sentence. The heuristic labels the sentences with the names of mission, instrument, re-analysis models, and science keywords taken from the Global Change Master Directory (GCMD) ontology. Additionally, it uses those labels to generate the dataset citation candidates. If the mission, instrument, and variable are sufficient to create the citation for the dataset the citation and the label the domain expert reviews the output without going through the NLP model. If the extracted label is not sufficient to label the dataset on its own, the sentence and its associated dataset labels will be inputted into the NER model. The model outputs the labeled sentence and the potential dataset candidates with their associated probabilities. The domain expert then reviews the NER model’s output and the correct labels are determined. The newly labelled papers can then be used as additional training data. This creates an iterative process for the approach to continuously improve. Because all the possible mentions are gathered by the model, the domain expert can quickly and easily label the papers resulting in large time savings.

Jacob Atkins↗

Knowledge Discovery for Early Failure Assessment of Complex Engineered Systems Using Natural Language Processing

Emerging complex engineered systems may have unexpected safety issues due to novel operational environments, increasing autonomy, human-machine interaction, and other factors. To prevent failures in operation or testing that necessitate costly redesign, it is desirable to predict likely failure modes early in the design process. Information about past engineering failures in natural language format presents one possible solution by enabling the retrieval of information that can inform new designs. However, identifying documents containing usable information and extracting the required information can be prohibitively time-consuming when implemented at scale. In this research, an automated natural language processing (NLP) framework is proposed to discover relevant knowledge from documents containing failure-related design information. The framework is applied to NASA’s Lessons Learned Information System (LLIS),which is publicly available. Documents containing usable information are filtered using two different NLP-based models. Next, from the identified usable documents, a failure taxonomy is extracted using a partitioned hierarchical topic modeling approach. Partitions of the document describe different sections of the failure taxonomy – i.e., failure, cause of failure, and recommendations – as indicated by the structure of the original document. The extracted failure taxonomy can be leveraged in early design failure assessment methods. Moreover, the framework can be used to identify documents containing usable failure-related design information from other databases and extract relevant information from these documents.

Documentation and Information Science↗

Knowledge Discovery for Early Failure Assessment of Complex Engineered Systems Using Natural Language Processing

Emerging complex engineered systems may have unexpected safety issues due to novel operational environments, increasing autonomy, human-machine interaction, and other factors. To prevent failures in operation or testing that necessitate costly redesign, it is desirable to predict likely failure modes early in the design process. Information about past engineering failures in natural language format presents one possible solution by enabling the retrieval of information that can inform new designs. However, identifying documents containing usable information and extracting the required information can be prohibitively time-consuming when implemented at scale. In this research, an automated natural language processing (NLP) framework is proposed to discover relevant knowledge from documents containing failure-related design information. The framework is applied to NASA’s Lessons Learned Information System (LLIS),which is publicly available. Documents containing usable information are filtered using two different NLP-based models. Next, from the identified usable documents, a failure taxonomy is extracted using a partitioned hierarchical topic modeling approach. Partitions of the document describe different sections of the failure taxonomy – i.e., failure, cause of failure, and recommendations – as indicated by the structure of the original document. The extracted failure taxonomy can be leveraged in early design failure assessment methods. Moreover, the framework can be used to identify documents containing usable failure-related design information from other databases and extract relevant information from these documents.

Documentation and Information Science↗

Search Enhancements using Natural Language Processing Techniques

NASA Goddard Earth Sciences Data and Information Services Center (GESDISC) is one of the 12 NASA Science Mission Directorate Data Centers. The main goal of GESDISC is to provide earth science data, information, and services to the earth science data community. Consequently, data discovery is at the center of our mission and our search engine is the primary tool for our users to interact, find, and access our data. Existing search approaches are largely focused on hard-matching of keywords in the search query with dataset metadata. Here we propose to expand the search by introducing a complementary natural language processing (NLP) search. At the heart of our proposed NLP search, we trained a joint embedding using scientific text corpus and a curated set of dataset metadata. The embedding learns the association between words in our dataset metadata and those of the scientific text corpus. This enables us to go beyond simple hard-matching of a query and data set metadata and have a notion of “similarity” between the search query and the datasets. We further integrated our NLP search into the Elastic Search (ES) framework leveraging similarity search capabilities offered through the “dense_vector” field type. Our preliminary evaluations show that our proposed NLP search has the potential to be utilized to complement the existing search engine and serve as a base for a dataset recommendation system.

Armin Mehrabian↗

A Brief Introduction to AI/ML Applications of Air Traffic Management Data at NASA Ames

This presentation will give a brief overview of several AI/ML This presentation will give a brief overview of several AI/ML projects that NASA Ames interns are exploring in partnership with NASA Aeronautic Research Institute (NARI) and the FAA. NASA is interested in Natural Language Processing (NLP) of various legacy text and speech data within air traffic management e.g., Notices To Airmen (NOTAMs), Letters of Agreement (LoAs), Standard Operating Procedures (SOPs), and Air Traffic Control Center audio briefings. Since our focus is on applying state of the art AI/ML tools to legacy air traffic management data, we first showcase the different data sources of interest followed by a brief introduction to the techniques and language models used. We present some exciting preliminary results on each topic including both unsupervised learning techniques (e.g., clustering) and other modern language models (e.g., BERT) that help extract useful information from these data sources that are interpretable by both man and machine.

Air Traffic Management↗

Wildfire Emergency Response Hazard Extraction and Analysis of Trends (HEAT) through Natural Language Processing and Time Series

Emerging wildfire operations aim to improve safety and performance through the integration of technologies including UAS and UTM. Recent advances in natural language processing (NLP) techniques, as well as the availability of wildfire incident reports, has made possible a large-scale analysis of wildfire hazards and trends. Identifying longitudinal trends will help us target risk mitigation and safety management activities. Note: This presentation does not include sound please disregard icon.

Sequoia R. Andrade↗

Verb Sense Disambiguation for Densifying Knowledge Graphs in Earth Science

We begin with an ambitious goal: to create a knowledge graph that spans the entire discipline of Earth science. In order to achieve this, we need to apply Natural Language Processing (NLP) techniques on Earth science journal articles to extract their semantic components for the graph. When sentences from Earth science journal articles are broken down into their semantic components and loaded onto a graph, the relationships among these semantic components are represented by the verbs in the sentences. However, since there are multiple verbs in English that can be used to denote the same meaning, the knowledge graph can become sparse and so can the results when we query the graph. In order to ensure quality results, it would be desirable to consolidate similar verbs into a single "class". So, this is the problem at hand: how do we make sure that multiple verbs that mean the same thing are represented as a single class of verb in the knowledge graph? Or in other words, how do we distinguish which meaning a particular verb takes given a particular sentence? In this poster, we demonstrate a potential technique to solve this problem.

Ashish Acharya↗

Integrating Human System Information with the Systems Platform for Aggregating and Relating Capabilities (SPARC)

Within Human Health and Performance, there exists a wealth of human system information that’s used a regular basis in support of NASA human exploration objectives, but the challenge is that all of this information was stored in multiple different locations and organized for specific uses, limiting its effectiveness and straining communication across multiple groups. To address this challenge, our project, the Systems Platform for Aggregating and Relating Capabilities (SPARC) was tasked with developing a new NASA internal web application that aggregates and relates multiple programs’ human system products, such as technical standards, program requirements and verifications, human system risks, research and evidence, and exploration capabilities, into one centralized platform that addresses the needs of human health and performance from multiple different perspectives. Using agile development methodologies, user-experience (UX) driven design principles, data visualization, and a strong emphasis on continuous improvement though consistent stakeholder engagement, the SPARC project released a beta version in less than 4 months, broadly launched version 1.0.0 Agency-wide three months after the beta, and has over 180 users in the first year of development. Our second year of development will see us moving from our initial capabilities to increasingly robust and complex integrations and visualizations, including Directed Acyclical Graphs (DAGs), natural language processing (NLP) for dynamic generation of relationships between the sources of truth, and an expansion into hierarchical levels of system design in support of the human exploration programs.

Data science↗

Transcribing Air Traffic Control System Command Center Planning Telecons Using Cloud-Based Automatic Speech Recognition

This paper addresses the challenge of using Automatic Speech Recognition (ASR) technology to transcribe regular teleconferences that happen between FAA Air Traffic Control System Command Center (ATCSCC) planners, stakeholders and air users. These planning teleconferences (aka telecons or planning webinars) are an integral part of managing air traffic in the U.S. National Airspace System (NAS). In particular, the meetings facilitate the creation and modification of various traffic management initiatives (TMIs), that are used to regulate the flow of air traffic. This is typically a human intensive process, requiring specialists to listen to the entire meeting audio (10-20 minutes duration) and inferring the state of the NAS (e.g., weather phenomenon) that was discussed. It would be advantageous to have digital transcripts of the audio and have useful information (e.g., related to TMIs) automatically extracted from the transcripts. In this regard, we are exploring the adoption of state-of-the-art speech to text and Natural Language Processing (NLP) tools that will achieve our objective of digitizing the webinar audio. Unfortunately, the highly technical phraseology present in the audio and limited data availability for model building make ASR difficult. To overcome this challenge, we have taken the critical first step in creating a human transcription dataset from ~20 hours of speech in the ATCSCC audio with the help of subject matter experts. A novelty of our work is the creation of a ground truth transcription dataset for ATCSCC teleconference webinars, which is particularly important for Aviation domain-specific NLP tasks. Using Microsoft Speech Studio, a cloud-based ASR platform, we have fine-tuned the English pre-trained ASR models (available in speech studio) and achieved an average word error rate (WER) of 6.81%. The baseline ASR also provides a digital version of each planning webinar, making it accessible and text-searchable for future references. Additionally, the transcriptions can serve as a bridge between raw audio data and a range of text-based NLP tasks, such as named entity recognition (NER) and intent classification, potentially enhancing the digital footprint of the webinars and other connected data sources. Our work has several potential applications. Firstly, the transcriptions can be analyzed to understand the complex decision process of creating, implementing and modifying TMIs and may also contribute to TMI prediction services. Secondly, our dataset and model can be used to develop more accurate ASR systems for aviation-specific language, which can bring about digital communication in the aviation industry (and aid current “voice only” communications, which are inherently error-prone). Lastly, the transcriptions themselves can be used as a valuable resource for training other NLP models.

Stephen S. B. Clarke↗