Engineering PapersSearch

SEARCH · Engineering Papers

Results for “NLP”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Nanolipoprotein particle (NLP) vaccine confers protection against Yersinia pestis aerosol challenge in a BALB/c mouse model

Introduction: Yersinia pestis is the etiological agent of plague, a disease that remains a concern as demonstrated by recent outbreaks in Madagascar. Infection with Y. pestis results in a rapidly progressing illness that can only be successfully treated with antibiotics given shortly after symptom onset. Live attenuated or whole cell inactivated vaccines confer protection against bubonic plague, but pneumonic plague has been more difficult to prevent. Novel effective subunit vaccine formulations may circumvent some of these shortfalls. Here, we compare the immunogenicity generated by an advanced subunit vaccine (F1V fusion protein) and a nanolipoprotein particle (NLP)-based vaccine. Methods: The NLP, a high-density lipoprotein mimetic, provides a nanoscale delivery platform for recombinant Y. pestis antigens LcrV (V) and F1. BALB/c mice were immunized via subcutaneous injection twice, three or four weeks apart. Four weeks later, splenocytes and sera were collected for immune profiling, and mice were challenged with aerosolized Y. pestis CO92. Results: Both formulations induced a strong IgG response against the F1 and V proteins, along with a robust memory B cell response and a balanced cell-mediated immune response as evidenced by both Th1- and Th2-related cytokines. The NLP-based vaccine induced a stronger cytokine response against F1, V, and F1V proteins relative to the F1V vaccine. As with F1V, the inclusion of Alhydrogel (Alu) in NLP vaccine formulations was critical for enhanced immunogenicity and protective efficacy. Mice that received two doses of F1:V:NLP + Alu and CpG were completely protected from a challenge with approximately eight median lethal doses of aerosolized Y. pestis CO92 and this protection confirmed the well-documented synergy between the F1 and V antigens in context of pneumonic plague. The NLPs have defined regions of polarity that facilitates the incorporation of a wide range of adjuvants and antigens with distinct physicochemical properties and are an excellent candidate platform for the development of multi-antigen vaccines.

F1

Natural Language Processing (NLP) Analysis of NOTAMs for Air Traffic Management Optimization

With new emerging technologies in the field of NLP, we explore their applications to digitize and analyze heritage Air Traffic Management (ATM) documents for planning and optimizing airspace operations. Specifically, this research focuses on harvesting semi-structured or un-structured information contained in Notices to Airmen (NOTAMs). Using NLP and other advanced data analytics, we will construct a data-driven framework which facilitates finding language patterns and the use of pretrained language models for classification and extraction of useful airspace constraints and restrictions. These may lead to tools that assist airspace users in understanding the constraints more efficiently, contributing to better route planning and safer execution. This paper explores three workflows entailing different NLP tasks. First, unsupervised techniques like word embedding and topic modeling are used for pattern finding and document classification. Second, a dataset is created by extracting information from the semi-structured NOTAM format as metadata for categorizing, visualizing, and extracting key entities driving NOTAM content. Third, modern pre-built deep learning based transformer models such as BERT, RoBERTa, and XLNet are evaluated on the question answering task, an even more robust approach to information extraction, as well as their respective fine-tuning tasks. In this work we include various performance metrics for the trained models to evaluate both accuracy and precision and we show that the models can be generalized for their respective tasks. The research work developed shows promise in uncovering trends in digital NOTAMs in the NAS and also offers a new framework for digitizing and inferring insights from free-form legacy NOTAMs, that are yet to be digitized. Video is an mp4 download, with a play time of 9 min 35 secs.

Natural Language Processing

Launch flexibility using NLP guidance and remote wind sensing

This paper examines the use of lidar wind measurements in the implementation of a guidance strategy for a nonlinear programming (NLP) launch guidance algorithm. The NLP algorithm uses B-spline command function representation for flexibility in the design of the guidance steering commands. Using this algorithm, the guidance system solves a two-point boundary value problem at each guidance update. The specification of different boundary value problems at each guidance update provides flexibility that can be used in the design of the guidance strategy. The algorithm can use lidar wind measurements for on pad guidance retargeting and for load limiting guidance steering commands. Examples presented in the paper use simulated wind updates to correct wind induced final orbit errors and to adjust the guidance steering commands to limit the product of the dynamic pressure and angle-of-attack for launch vehicle load alleviation.

Cramer, Evin J.

Informing NLP Learning Tasks by Tracking User Features: An ASRS Use Case using Kaona

There has been growing interest in utilizing natural language processing (NLP) algorithms in Aviation Safety. This interest has extended to leveraging the decades of records publicly available on the Aviation Safety Reporting System (ASRS). While related literature has given more emphasis in lessons learned from the narratives, our prior work has focused on using NLP to support narrative search in the ASRS. Specifically, we evaluated if the use of alternative search mechanisms to keyword search, such as the retrieval of related narratives even without matching keywords could improve narrative discovery. A difficulty in experimenting alternative search mechanisms in any information retrieval task is the lack of ground truth. To address this limitation, we propose Kaona, a lightweight interface which enables the prototyping of alternative search retrieval tasks, by tracking user experience both explicitly (user-specified feedback), or implicitly (user navigation through interface affordances). Differently from distracting requests for feedback during user navigation, Kaona collects explicit feedback from users by mapping them to affordances which support the user workflow, while obtaining ground truth information for learning tasks.

human-computer-interaction

Harmonizing Multi-Disciplinary Data for Applied Research: A Comprehensive Analysis of GES DISC Datasets through NLP and TF-IDF Techniques

Data centers distribute data encompassing multiple disciplines, making it necessary to evaluate dataset applicability to applied research. Typically, these datasets are generated by specialized science teams and consist of single-discipline data, such as atmospheric temperature, pressure, and precipitation, leading to unique dataset formats and access services. However, in applied research, the utilization of datasets from multiple disciplines is commonly necessary. The increasing availability of research literature citing Earth Science datasets presents an opportunity to analyze the usage of datasets in multi-disciplinary research. This study proposes a novel approach wherein research publications citing datasets archived at the GES DISC (Goddard Earth Sciences Data and Information Services Center) are collected, and each publication is associated with specific research topics through the application of Natural Language Processing (NLP), using the Term Frequency Inverse Document Frequency (TF-IDF) technique on the publication titles and abstracts. Through this analysis, we gain insights into the distribution of dataset disciplines as they are being used in various applied research areas. This knowledge is essential for the development of dataset tools and services tailored to effectively support applied research studies, as it enhances a data center's comprehension of how datasets from multiple disciplines are integrated into research endeavors.

Infometrics

Informing Plant Asset Reliability and Availability Through AI-Driven Analysis of Operator Logs

The availability and reliability of nuclear power plant (NPP) structures, systems, and components (SSCs) are critical parameters for NPP safety. Tracking these parameters is necessary but costly and labor-intensive, requiring the collection and evaluation of SSC event data such as shutdowns, startups, and failures. To show how these events are needed for the parameters an example is given: one measure of reliability is based on the number of equipment failure events and the number of run hours (i.e., the time from a startup event to a shutdown event). Here, this work investigates using artificial intelligence (AI) to mine NPP operator log entry texts for SSC event data. Four AI approaches were explored for identifying these events, including natural language processing (NLP) methods, generative AI, generative AI combined with NLP, and topic modeling. A key challenge addressed with all four approaches is the brevity of operator log entries. Among these four a neural network–based NLP method was shown to be the most promising for this application, achieving F1 scores of 86.0% for shutdowns, 92.2% for startups, and 80.4% for failures on a subject-matter-expert-curated dataset from NPP operator logs, compared to a baseline of 66.6% for a random classifier. This shows that NLP methods can perform better than generative AI. Additionally, the NLP methods combined with generative AI were shown to perform better than generative AI alone. Generative AI was most successful at providing the background information for the NLP methods to use. This work demonstrates the potential to use AI to automate parameter collection from NPP operator log entries and other records.

97 - MATHEMATICS AND COMPUTING

Excipient screening by lyophilization provides insights into spray drying formulations for nanoparticle vaccines

Nanoparticles have shown great promise as delivery platforms in the development of tunable and safe vaccines. Nanolipoprotein particles (NLPs), also known as nanodiscs, are discoidal nanoparticles composed of a lipid bilayer stabilized at their periphery by apolipoproteins. Under the right conditions, the NLP self-assembly process is highly customizable in terms of lipids and apolipoprotein constituents, allowing for tunable physical and chemical characteristics. This flexibility allows a wide range of vaccine antigens and adjuvants to be incorporated onto the NLP platform for tailored vaccine design. The stability of NLPs during long term storage is a very important factor in developing a vaccine delivery platform suitable for widespread global use. When stored in a solution for extended periods of time, NLPs dissociate into their corresponding lipids and protein constituents, leading to particle degradation. Proper stabilization of NLPs can often be achieved by lyophilization (i.e. freeze-drying), a method widely used for various applications including pharmaceuticals. This process, however, can be damaging to particles without the presence of lyoprotectants or excipients that help maintain particle stability during lyophilization. Another method used to stabilize vaccines and pharmaceuticals is spray drying, a process that converts liquid formulations into dry powders through controlled heating and airflow. While spray drying is rapid, scalable, and cost-effective, lyophilization is typically a gentler process that better retains biomolecule structure and function. Both processes use excipients for particle stabilization, so lyophilization can be used as a surrogate to test stability of NLPs, to down-select formulations that may withstand the harsher conditions of spray drying. To screen different formulations, NLPs were synthesized and purified to homogeneity by size exclusion chromatography (SEC) and samples were prepared with a wide range of lyoprotectants and/or excipients. To assess the protective effects of excipients on NLPs upon spray drying, both pre- and post-lyophilized samples were analyzed by SEC. To assess the protective effects upon heating (encountered during the spray drying process), NLP samples were incubated at elevated temperatures prior to SEC analysis. The lyoprotectants and excipients evaluated in this study had different efficiencies in protecting NLPs during lyophilization and heating tests. Trehalose, for example, exhibits stabilization on NLPs both upon lyophilization and heating whereas leucine accelerated NLP dissociation. Although some lyoprotectants are effective by themselves, different combinations can decrease the stabilization of NLPs. Shelf-stable vaccines that do not require cold-chain storage are essential for global accessibility and our findings provide fundamental insight into how to advance NLP-based vaccines for these applications.

Serrano, Litzay J [Lawrence Livermore National Lab

ARCH: Large-scale knowledge graph via aggregated narrative codified health records analysis

Objective: Electronic health record (EHR) systems contain a wealth of clinical data stored as both codified data and free-text narrative notes (NLP). The complexity of EHR presents challenges in feature representation, information extraction, and uncertainty quantification. Here, to address these challenges, we proposed an efficient Aggregated naRrative Codified Health (ARCH) records analysis to generate a large-scale knowledge graph (KG) for a comprehensive set of EHR codified and narrative features. Methods: Using data from 12.5 million Veterans Affairs patients, ARCH first derives embedding vectors and generates similarities along with associated p-values to measure the strength of relatedness between clinical features with statistical certainty quantification. Next, ARCH performs a sparse embedding regression to remove indirect linkage between features to build a sparse KG. Finally, ARCH was validated on various clinical tasks, including detecting known relationships between entity pairs, predicting drug side effects, disease phenotyping, as well as sub-typing Alzheimer’s disease patients. Results: ARCH produces high-quality clinical embeddings and KG for over 60,000 codified and narrative EHR concepts. The KG and embeddings are visualized in the R-shiny powered web-API.3 ARCH achieved high accuracy in detecting EHR concept relationships, with AUCs of 0.926 (codified) and 0.861 (NLP) for similar EHR concepts, and 0.810 (codified) and 0.843 (NLP) for related pairs. It detected drug side effects with a 0.723 AUC, which improved to 0.826 after fine-tuning. Using both codified and NLP features, the detection power increased significantly. Compared to other methods, ARCH has superior accuracy and enhances weakly supervised phenotyping algorithms’ performance. Notably, it successfully categorized Alzheimer’s patients into two subgroups with varying mortality rates. Conclusion: The proposed ARCH algorithm generates large-scale high-quality semantic representations and knowledge graph for both codified and NLP EHR features, useful for a wide range of predictive modeling tasks.

Electronic health records

Language models for materials discovery and sustainability: Progress, challenges, and opportunities

Significant advancements have been made in one of the most critical branches of artificial intelligence: natural language processing (NLP). These advancements are exemplified by the remarkable success of OpenAI’s GPT-3.5/4 and the recent release of GPT-4.5, which have sparked a global surge of interest akin to an NLP gold rush. Here, in this article, we offer our perspective on the development and application of NLP and large language models (LLMs) in materials science. We begin by presenting an overview of recent advancements in NLP within the broader scientific landscape, with a particular focus on their relevance to materials science. Next, we examine how NLP can facilitate the understanding and design of novel materials and its potential integration with other methodologies. To highlight key challenges and opportunities, we delve into three specific topics: (i) the limitations of LLMs and their implications for materials science applications, (ii) the creation of a fully automated materials discovery pipeline, and (iii) the potential of GPT-like tools to synthesize existing knowledge and aid in the design of sustainable materials.

36 MATERIALS SCIENCE

Protein–Protein Interaction Networks Derived from Classical and Machine Learning-Based Natural Language Processing Tools

The study of protein-protein interactions (PPIs) provides insight into various biological mechanisms, including the binding of antibodies to antigens, enzymes to inhibitors or promoters, and receptors to ligands. Recent studies of PPIs have led to significant biological breakthroughs. For example, the study of PPIs involved in the human:SARS-CoV-2 viral infection mechanism aided in the development of the SARS-CoV-2 vaccines. Though several databases exist for the manual curation of PPI networks, text mining methods have been routinely demonstrated as useful alternatives for newly studied or understudied species where databases are incomplete. Here, the relationship extraction (RE) performance of several open-source classical text processing, machine learning (ML)-based natural language processing (NLP), and large language model (LLM)-based NLP tools were compared. Overall, our results indicated that networks derived from classical methods tend to have high true positive rates at the expense of having overconnected-networks, ML-based NLP methods have lower true positive rates but networks with the closest structures to the target network, and LLM-based NLP methods tend to exist in-between the two other approaches, with variable performances. Finally, the selection of a specific NLP approach should be tied to the needs of a study and text availability, as models varied in performance due to the amount of text provided.

59 BASIC BIOLOGICAL SCIENCES

A Model Based Approach to Extract Health Information from Textual Data

In current nuclear power plants (NPPs) a large amount of condition-based data is being generated and stored to assess and monitor component health and performance. The format of this data can be either numeric (e.g., pump vibration data) or textual (e.g., condition report which assess component health). While assessing component health from numeric data can be performed with a large variety of methods, the extraction of information from textual data still remains a challenge. Natural language processing (NLP) methods are starting to be deployed in current NPPs mainly to filter out incident reports (IRs) that are not safety related by employing supervised machine learning methods. However, these methods do not really provide the quantitative information that might be contained in IRs. This paper presents an approach to extract information from textual data (e.g., from IRs, maintenance reports) that is based on NLP data analytics methods coupled with model-based system engineer (MBSE) models. NLP methods are employed to perform syntactic and semantic analyses. Syntactic analysis analyzes the grammatical structure of a sentence; such analysis includes: part of speech (POS) tagging (i.e., identification of grammatic elements of each string - e.g., nouns, verbs), named entity recognition (i.e., identification of text entities - e.g., names, dates, events), and relation extraction (e.g., coreference resolution). On the other hand, semantic analysis is designed to analyze the logic structure of a sentence. Through a specific set of rules, our methods can identify whether a sentence contains health information of a component (e.g., degraded performance, anomaly behavior) or the causal relationship between two events (i.e., a cause-effect pair). An innovative element of our approach is that semantic analysis relies on MBSE models to identify links between textual elements. MBSE are diagrams designed to represent system and component dependencies (from both a form and functional point of view). In our approach, MBSE models emulate system engineer knowledge about component/system architecture. This paper presents in detail how the integration of NLP methods and MBSE models is performed. Few analysis examples focusing on centrifugal pumps are presented.

97 - MATHEMATICS AND COMPUTING

Molecular property prediction for very large databases with natural language processing: a case study in ionic liquid design

The prospect of using artificial intelligence (AI) to accurately screen very large databases of compounds for multiple properties has yet to be realized. Here, we explore this possibility using ionic liquids (ILs) which offer unique physicochemical properties and excellent tunability, making them highly versatile solvents for various research applications. Screening millions of potential ILs for the best perfomance for use in specific tasks with experimental methods alone however, is impractical. Further, traditional’ physics-based computational chemistry is hindered by high computational cost. To address this challenge, we leverage a natural language processing (NLP)-based molecular embedding technique with advanced machine learning (ML) models to predict seven key IL properties: viscosity, density, ionic conductivity, surface tension, melting temperature, toxicity, and water solubility. Comprehensive datasets for these properties are obtained, then NLP featurization with Mol2vec is compared with other featurization techniques such as 2D Morgan fingerprints, and 3D quantum chemistry-derived sigma profiles. NLP-based featurization exhibited the best predictive performance, achieving the highest R 2 and lowest RMSE values for all the studied IL properties. Further, we present case studies of how ILs might be screened using combined property criteria for practical cases – lignocellulosic biomass processing, CO 2 capture, and optimal electrolytes for batteries – screening a novel database of ∼10.6 million generated feasible ILs. The results introduce NLP as a powerful tool for engineering many designer solvents with desirable properties for task specific applications.

Mohan, Mood [Oak Ridge National Laboratory (ORNL),

Cell-Free Screening, Production and Animal Testing of a STI-Related Chlamydial Major Outer Membrane Protein Supported in Nanolipoproteins

Background: Vaccine development against Chlamydia, a prevalent sexually transmied infection (STI), is imperative due to its global public health impact. However, significant challenges arise in the production of effective subunit vaccines based on recombinant protein antigens, particularly with membrane proteins like the Major Outer Membrane Protein (MOMP). Methods: Cellfree protein synthesis (CFPS) technology is an aractive approach to address these challenges as a method of high-throughput membrane protein and protein complex production coupled with nanolipoprotein particles (NLPs). NLPs provide a supporting scaffold while allowing easy adjuvant addition during formulation. Over the last decade, we have been working toward the production and characterization of MOMP-NLP complexes for vaccine testing. Results: The work presented here highlights the expression and biophysical analyses, including transmission electron microscopy (TEM) and dynamic light scaering (DLS), which confirm the formation and functionality of MOMP-NLP complexes for use in animal studies. Moreover, immunization studies in preclinical models compare the past and present protective efficacy of MOMP-NLP formulations, particularly when co-adjuvanted with CpG and FSL1. Conclusion: Ex vivo assessments further highlight the immunomodulatory effects of MOMP-NLP vaccinations, emphasizing their potential to elicitrobust immune responses. However, further research is warranted to optimize vaccine formulations further, validate efficacy against Chlamydia trachomatis, and beer understand the underlying mechanisms of immune response.

60 APPLIED LIFE SCIENCES

Internship Abstract and Final Reflection

The primary objective for this internship is the evaluation of an embedded natural language processor (NLP) as a way to introduce voice control into future space suits. An embedded natural language processor would provide an astronaut hands-free control for making adjustments to the environment of the space suit and checking status of consumables procedures and navigation. Additionally, the use of an embedded NLP could potentially reduce crew fatigue, increase the crewmember's situational awareness during extravehicular activity (EVA) and improve the ability to focus on mission critical details. The use of an embedded NLP may be valuable for other human spaceflight applications desiring hands-free control as well. An embedded NLP is unique because it is a small device that performs language tasks, including speech recognition, which normally require powerful processors. The dedicated device could perform speech recognition locally with a smaller form-factor and lower power consumption than traditional methods.

Sandor, Edward