Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “natural language processing (NLP)”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Molecular property prediction for very large databases with natural language processing: a case study in ionic liquid design

The prospect of using artificial intelligence (AI) to accurately screen very large databases of compounds for multiple properties has yet to be realized. Here, we explore this possibility using ionic liquids (ILs) which offer unique physicochemical properties and excellent tunability, making them highly versatile solvents for various research applications. Screening millions of potential ILs for the best perfomance for use in specific tasks with experimental methods alone however, is impractical. Further, traditional’ physics-based computational chemistry is hindered by high computational cost. To address this challenge, we leverage a natural language processing (NLP)-based molecular embedding technique with advanced machine learning (ML) models to predict seven key IL properties: viscosity, density, ionic conductivity, surface tension, melting temperature, toxicity, and water solubility. Comprehensive datasets for these properties are obtained, then NLP featurization with Mol2vec is compared with other featurization techniques such as 2D Morgan fingerprints, and 3D quantum chemistry-derived sigma profiles. NLP-based featurization exhibited the best predictive performance, achieving the highest R 2 and lowest RMSE values for all the studied IL properties. Further, we present case studies of how ILs might be screened using combined property criteria for practical cases – lignocellulosic biomass processing, CO 2 capture, and optimal electrolytes for batteries – screening a novel database of ∼10.6 million generated feasible ILs. The results introduce NLP as a powerful tool for engineering many designer solvents with desirable properties for task specific applications.

Mohan, Mood [Oak Ridge National Laboratory (ORNL),↗

PhysBERT: A text embedding model for physics scientific literature

The specialized language and complex concepts in physics pose significant challenges for information extraction through Natural Language Processing (NLP). Central to effective NLP applications is the text embedding model, which converts text into dense vector representations for efficient information retrieval and semantic analysis. In this work, we introduce PhysBERT, the first physics-specific text embedding model. Pre-trained on a curated corpus of 1.2 × 106 arXiv physics papers and fine-tuned with supervised data, PhysBERT outperforms leading general-purpose models on physics-specific tasks, including the effectiveness in fine-tuning for specific physics subdomains.

Hellert, Thorsten (ORCID:0000000227970926)↗

EXFOR-NSR PDF database: a system for nuclear knowledge preservation and data curation

Current needs of nuclear science and technology include complete, well-documented, and easily verifiable nuclear data. The complete data records require supporting nuclear bibliography, presently stored in dedicated libraries, in addition, to actual data. Additionally, experimental nuclear reaction data (EXFOR) and Nuclear Science References (NSR) databases contain compilations based on primary (journals) and secondary (conference proceedings, theses, preprints, etc.) publications, and data received from authors via private communications. The secondary library materials and private communications often represent a bottleneck for nuclear data verification, compilation, evaluation, and dissemination activities. To address this issue, bibliographic materials were scanned into PDF (Portable Document Format) files and uploaded in a relational database. The traditional scope of nuclear databases that includes meta-data and numbers derived from data in specialized formats was broadened to accommodate the large volumes of original nuclear data publications. The complete PDF publication files were stored in a relational database as Binary Large OBjects (BLOB). This unique collection of nuclear data compilations and supporting publications generate many opportunities for machine learning applications. The Web interfaces for authorized and public access to the EXFOR-NSR nuclear publications database were implemented at the U.S. National Nuclear Data Center, https://www.nndc.bnl.gov/ and IAEA Nuclear Data Section, https://www-nds.iaea.org/ . The current system is complementary to major nuclear libraries and narrowly focused on nuclear data compilation and evaluation procedures. The contents of the PDF database, details of implementation, and Web interface are described. New capabilities for data curation, knowledge preservation, worldwide dissemination, and natural language processing (NLP) applications are given.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Attention-based quantum tomography

Abstract With rapid progress across platforms for quantum systems, the problem of many-body quantum state reconstruction for noisy quantum states becomes an important challenge. There has been a growing interest in approaching the problem of quantum state reconstruction using generative neural network models. Here we propose the ‘attention-based quantum tomography’ (AQT), a quantum state reconstruction using an attention mechanism-based generative network that learns the mixed state density matrix of a noisy quantum state. AQT is based on the model proposed in ‘Attention is all you need’ by Vaswani et al (2017 NIPS ) that is designed to learn long-range correlations in natural language sentences and thereby outperform previous natural language processing (NLP) models. We demonstrate not only that AQT outperforms earlier neural-network-based quantum state reconstruction on identical tasks but that AQT can accurately reconstruct the density matrix associated with a noisy quantum state experimentally realized in an IBMQ quantum computer. We speculate the success of the AQT stems from its ability to model quantum entanglement across the entire quantum system much as the attention model for NLP captures the correlations among words in a sentence.

97 MATHEMATICS AND COMPUTING↗

Radio galaxy zoo EMU: towards a semantic radio galaxy morphology taxonomy

We present a novel natural language processing (NLP) approach to deriving plain English descriptors for science cases otherwise restricted by obfuscating technical terminology. We address the limitations of common radio galaxy morphology classifications by applying this approach. We experimentally derive a set of semantic tags for the Radio Galaxy Zoo EMU (Evolutionary Map of the Universe) project and the wider astronomical community. We collect 8486 plain English annotations of radio galaxy morphology, from which we derive a taxonomy of tags. The tags are plain English. The result is an extensible framework, which is more flexible, more easily communicated, and more sensitive to rare feature combinations, which are indescribable using the current framework of radio astronomy classifications.

79 ASTRONOMY AND ASTROPHYSICS↗

Domain-specific text embedding model for accelerator physics

Accelerator physics presents unique challenges for natural language processing (NLP) due to its specialized terminology and complex concepts. A key component in overcoming these challenges is the development of robust text embedding models that transform textual data into dense vector representations, facilitating efficient information retrieval and semantic understanding. In this work, we introduce AccPhysBERT, a sentence embedding model fine-tuned specifically for accelerator physics. Our model demonstrates superior performance across a range of downstream NLP tasks, surpassing existing models in capturing the domain-specific nuances of the field. We further showcase its practical applications, including semantic paper-reviewer matching and integration into retrieval-augmented generation systems, highlighting its potential to enhance information retrieval and knowledge discovery in accelerator physics. Published by the American Physical Society 2025

Hellert, Thorsten (ORCID:0000000227970926)↗

Towards Automatic Mapping of Vulnerabilities to Attack Patterns using Large Language Models

With the advent of new devices and applications, cyber attack surface is continuously evolving due to the emergence of new attack techniques and vulnerabilities. Hence, security management tool must assess the cyber risk of an enterprise at regular interval basis through comprehensively identifying associations among attack techniques, weakness, and vulnerabilities. However, existing repositories providing such associations are incomplete (i.e., missing associations), inducing the likelihood of undermining the risk of particular set of attack techniques. Moreover, such associations still rely on manual interpretation, which is slow compared to attack speed and ineffective for the increasing list of vulnerabilities and attack actions. Therefore, there is an urge to develop methodologies for automatically associating vulnerabilities to all relevant attack techniques. In this paper, we present a framework, named VWC-MAP, that can automatically identify all relevant attack techniques of a vulnerability via weakness based on their text descriptions, applying natural language process (NLP) techniques. To achieve that, we present a novel two-tiered classification approach, where the first tier classifies vulnerabilities to weakness, and the second tier classifies weakness to attack techniques. This research has improved the scalability of the current state-of-the-art tool to make vulnerability to weakness mapping significantly faster. Moreover, this paper presents two novel approaches for weakness to attack technique mapping applying Text-to-Text and link prediction techniques. Our experiment results cross-validated through cyber-security experts show that VWC-MAP can associate vulnerabilities to weakness types with 87% accuracy and to new attack patterns with 80% accuracy.

Das, Siddhartha Shankar↗

Characterizing Quantum Classifier Utility in Natural Language Processing Workflows

Quantum Natural Language Processing (QNLP) develops natural language processing (NLP) models for deployment on quantum computers. We explore feature and data prototype selection techniques to address challenges posed by encoding high dimensional features. Our study builds quantum circuit classifiers that includes classical feature pre-processing, quantum embedding and quantum model training. The quantum models are built on 4 or 6 qubits and the quantum neural network (QNN) uses the established bricklayer design. We compare the dependence of model performance (in terms of accuracy and F1 scores) on feature length, embedding gates and parameterized unitary design. We compare the performance of quantum machine learning models to classical convolution neural network model (CNN) on binary and multi-class classification tasks using two datasets of synthetic features and labels. The first is the ECP-CANDLE P3B3 dataset a corpus of synthetically generated cancer pathology reports. The second dataset is extracted from well-known benchmark dataset (MADELON) - features are generated with a combination of informative, repeated and uninformative features. Both datasets are used for binary classification and multi-class classification with 3 classes. We observe robust, accurate performance from all models on the binary classification tasks, but multiclass classification is a challenge for the quantum models-there is a notable decrease in accuracy when using 3 classes. Overall the performance is comparable in terms of recall and accuracy between QNNs and CNNs, even with large datasets. These results provide a point of comparison between quantum and classical models on real-world datasets.

Hamilton, Kathleen↗

Codon2Vec v1.0

Background: Codon2Vec is an embedding neural network that predicts 'high' or 'low' gene expression directly from the protein-coding sequences. Embedding neural networks are commonly used for natural language processing (NLP) applications. Analogous to how an English sentence is a string of words, a gene can be thought of as a string of codons. Similar to how NLP neural networks model English sentences as a non-random sequence of words, we considered a coding sequence as a non-random non-overlapping array of codons (k-mers of length = 3). Value Proposition: - Codon2Vec achieved a high median AUC-ROC score of 83.8% when trained and applied to transcriptomic data from 300 fungal species - Unlike Codo2Vec, conventional methods predicting for expression based on codon usage rely on a priori knowledge of optimal codons or a set of reference genes. - Unlike Codon2vec, these methods do not account for the effect of codon order on gene expression. - Codon2Vec neural network bypasses the need for artisanal feature selection step that is necessary for traditional machine learning models.

Wint, Rhondene↗

labquake_future_prediction

The labquake_future_prediction code is a collection of python modules and scripts that serves as supporting information for the article “Predicting future laboratory fault friction through deep learning” for publication in the journal of “Geophysical Research Letters”. It is designed to predict laboratory fault slips in the immediate future by scanning continuous acoustic emission (AE) waveforms recorded in laboratory biaxial shear experiments. The predictions are made with a deep learning model based on convolutional encoder-decoder (CED) models and the Transformer model primarily developed for Natural Language Processing (NLP). The deep learning model is trained with the tensorflow package using publicly available laboratory data sets in standard binary file format in numpy. The utility functions for reading data files, configuring model hyperparameters, constructing the CED and Transformer models, training and testing of the models are defined in python module files. The workflow of training the models for labquake future predictions and the multiple GPU’s rapid model hyperparameter optimization as described in the journal article, are demonstrated in accompanying python script files and Jupyter notebooks.

Wang, Kun↗

FrESCO

The National Cancer Institute (NCI) monitors population level cancer trends as part of its Surveillance, Epidemiology, and End Results (SEER) program. This program consists of state or regional level cancer registries which collect, analyze, and annotate cancer pathology reports. From these annotated pathology reports, each individual registry aggregates cancer phenotype information and summary statistics about cancer prevalence to facilitate population level monitoring of cancer incidence. Extracting cancer phenotype from these reports is a labor intensive task, requiring specialized knowledge about the reports and cancer. Automating this information extraction process from cancer pathology reports has the potential to improve not only the quality of the data by extracting information in a consistent manner across registries, but to improve the quality of patient outcomes by reducing the time to assimilate new data and enabling time-sensitive applications such as precision medicine. Here we present FrESCO: Framework for Exploring Scalable Computational Oncology, a modular deep-learning natural language processing (NLP) library for extracting pathology information from clinical text documents.

Spannaus, Adam [Oak Ridge National Lab. (ORNL), Oa↗

Coreii - Scout

COREII Scout employs React, Vite, TypeScript, Tailwind, and Daisy UI for its graphical user interface (GUI), offering both dark and light modes. The code is modular, with components and reusable wrappers to enhance efficiency. The primary goal of COREII Scout is to aid analysts in collecting and analyzing various sources related to cyber attacks, utilizing models to automate the report writing process. It uses Named Entity Recognition (NER), a type of Natural Language Processing (NLP), to extract key entities from each source. Analysts review and classify these entities using the COREII Attack Chain Estimator (ACE), adding their comments. Ultimately, a Large Language Model (LLM) generates a detailed report with user guidance. This setup ensures a streamlined and effective approach to cyber attack analysis and reporting.

Pluth, Adam [Idaho National Laboratory (INL), Idah↗

A meta-analysis of semantic classification of citations

The aim of this literature review is to examine the current state of the art in the area of citation classification. In particular, we investigate the approaches for characterizing citations based on their semantic type. We conduct this literature review as a meta-analysis covering 60 scholarly articles in this domain. Although we included some of the manual pioneering works in this review, more emphasis is placed on the later automated methods, which use Machine Learning and Natural Language Processing (NLP) for analyzing the fine-grained linguistic features in the surrounding text of citations. The sections are organized based on the steps involved in the pipeline for citation classification. Specifically, we explore the existing classification schemes, data sets, pre-processing methods, extraction of contextual and non-contextual features, and the different types of classifiers and evaluation approaches. The review highlights the importance of identifying the citation types for research evaluation, the challenges faced by the researchers in the process, and the existing research gaps in this field.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Unstructured clinical notes within the 24 hours since admission predict short, mid & long-term mortality in adult ICU patients

Mortality prediction for intensive care unit (ICU) patients is crucial for improving outcomes and efficient utilization of resources. Accessibility of electronic health records (EHR) has enabled data-driven predictive modeling using machine learning. However, very few studies rely solely on unstructured clinical notes from the EHR for mortality prediction. In this work, we propose a framework to predict short, mid, and long-term mortality in adult ICU patients using unstructured clinical notes from the MIMIC III database, natural language processing (NLP), and machine learning (ML) models. Depending on the statistical description of the patients’ length of stay, we define the short-term as 48-hour and 4-day period, the mid-term as 7-day and 10-day period, and the long-term as 15-day and 30-day period after admission. We found that by only using clinical notes within the 24 hours of admission, our framework can achieve a high area under the receiver operating characteristics (AU-ROC) score for short, mid and long-term mortality prediction tasks. The test AU-ROC scores are 0.87, 0.83, 0.83, 0.82, 0.82, and 0.82 for 48-hour, 4-day, 7-day, 10-day, 15-day, and 30-day period mortality prediction, respectively. We also provide a comparative study among three types of feature extraction techniques from NLP: frequency-based technique, fixed embedding-based technique, and dynamic embedding-based technique. Lastly, we provide an interpretation of the NLP-based predictive models using feature-importance scores.

60 APPLIED LIFE SCIENCES↗

COVIDScholar: An automated COVID-19 research aggregation and analysis platform

The ongoing COVID-19 pandemic produced far-reaching effects throughout society, and science is no exception. The scale, speed, and breadth of the scientific community’s COVID-19 response lead to the emergence of new research at the remarkable rate of more than 250 papers published per day. This posed a challenge for the scientific community as traditional methods of engagement with the literature were strained by the volume of new research being produced. Meanwhile, the urgency of response lead to an increasingly prominent role for preprint servers and a diffusion of relevant research through many channels simultaneously. These factors created a need for new tools to change the way scientific literature is organized and found by researchers. With this challenge in mind, we present an overview of COVIDScholar https://covidscholar.org/, an automated knowledge portal which utilizes natural language processing (NLP) that was built to meet these urgent needs. The search interface for this corpus of more than 260,000 research articles, patents, and clinical trials served more than 33,000 users at an average of 2,000 monthly active users and a peak of more than 8,600 weekly active users in the summer of 2020. Additionally, we include an analysis of trends in COVID-19 research over the course of the pandemic with a particular focus on the first 10 months, which represents a unique period of rapid worldwide shift in scientific attention.

60 APPLIED LIFE SCIENCES↗

The Carbon Storage Technical Viability Approach (CS TVA)

The Carbon Storage Technical Viability Approach (CS TVA) StoryMap provides an in-depth overview of the products created during the CS TVA research effort. In detail, the StoryMap addresses the CS TVA Matrix, Database Version 2.0, Database Catalog, Workflow, and Data Availability Result Database, discussing how each was developed and implemented. Information on how the matrix, database, and database catalog are interconnected, and their usage is also explained. The workflow section provides information on the CS TVA product development from the data-gathering stage to the final data availability results. A section on an expansion of the CS TVA workflow that utilizes Natural Language processing (NLP) section was included. Finally, the Data Availability Results Database is discussed. These results provide data science-informed insights into potential data gaps when assessing the viability of carbon storage in a given area or region.

Carbon Storage↗

Retaining Systems Engineering Model Meaning Through Transformation: Demo 2

Digital engineering strategies typically assume that digital engineering models interoperate seamlessly across the multiple different engineering modeling software applications involved, such as model- based systems engineering (MBSE), mechanical computer-aided design (MCAD), electrical computer-aided design (ECAD), and other engineering modeling applications. The presumption is that the data schema in these modeling software applications are structured in the familiar flat- tabular schema like any other software application. Engineering domain-specific applications (e.g., systems, mechanical, electrical, simulation) are typically designed to solve domain-specific problems, necessarily excluding explicit representations of non-domain information to help the engineer focus on the domain problems (system definition, design, simulation). Such exclusions become problematic in inter-domain information exchange. The obvious assumptions of one domain might not be so obvious to experts in another domain. Ambiguity in domain-specific language can erode the ability to enable different domain modeling applications to interoperate, unless the underlying language is understood and used as the basis for translation from one application to another. The engineering modeling software application industry has struggled for decades to enable these applications to interoperate. Industry standards have been developed, but they have not unified the industry. Why is this? The authors assert that the industry has relied on traditional database integration methods. The basic issue prohibiting successful application integration then is that traditional database-driven integration does not consider the distinct languages of each domain. An engineering models meaning is expressed through the underlying language of that engineering domain. In essence, traditional integration methods do not retain the semantic context (meaning) of the model. The basis of this research stems from the widely held assumption that systems engineering models are (or can be) structured according to the underlying semantic ontology of the model. This assumption can be imagined from two thoughts. 1) Digital systems engineering models are often represented using graph theory (the graph of a complex systems model can contain millions of nodes and edges). When examining the nodes one at a time and following the outbound edges of each node one by one, one can end up with rudimentary statements about the model (i.e., node A relates to node B), as in a semantic graph. 2) Likewise, from the study of natural languages, a sentence can be structured into unambiguous triples of subject-predicate-object within formal and highly expressive semantic ontologies. The rudimentary statements about a systems model discerned with graph theory closely mimic the triples used in the ontologies that try to structure natural languages. In other words, a systems models semantic graph can be (or is) structured into an ontology. Additionally, it is well established in industry that through natural language processing (NLP), which provides the means to create language structures, that computers can interpret ontological graphs. Therefore, the authors hypothesized that if the integrity of the underlying semantic structure of a systems model is retained, the contextual meaning of the model is retained. By structuring system models into the triples of the underlying ontology during the transformation from one MBSE application to another, the authors have provided a proof of the concept that the meaning of a system model can be retained during transformation. The authors assert that this is the missing ingredient in effective systems model-to-model interoperability. ACKNOWLEDGEMENTS The authors would like to thank the FY19 Model Interoperability team members who provided a solid foundation for the FY20 team to leverage: John McCloud, for the work he did to guide us toward the right use of technology that will appropriately discover and manipulate ontologies. Carlos Tafoya, for the work he did to develop an application programming interface (API)/Adapter that would export ontology-based data from GENESYS. Peter Chandler, for the work he did to architect our overall integration solution, with an eye toward the future that would influence a large-scale federated production-level systems engineering digital model ecosystem.

42 ENGINEERING↗

MVP-CTL Fellowship for Suicide Intervention (CRADA Final Report)

Research conducted with Crisis Text Line (CTL) served as an important stepping stone in the group’s ongoing work of developing natural language processing (NLP) methodologies to discover indicators of suicide factors in the U.S. veteran population. With rising cases of suicide, detecting risk factors among veterans is of growing concern. Socio-economic factors are often under-reported along structured medical variables, and extracting information through NLP methods may improve the sensitivity of suicide prediction models. The research explored through the CTL collaboration provided insights on suicide crisis detection and de-escalation, experience in applying scalable data analysis on real world datasets, contributions to study designs for detecting rare-event and produced transferable methods applicable to the domain of veteran suicide mitigation. Tools developed during this research directly benefits ongoing work on the Million Veteran Project (MVP) collaboration with the Department of Veteran Affairs (VA).

99 GENERAL AND MISCELLANEOUS↗