Engineering PapersSearch

SEARCH · Engineering Papers

Results for “text extraction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Rapid Adaptation of Chemical Named Entity Recognition Using Few-Shot Learning and LLM Distillation

Named entity recognition (NER) has been widely used in chemical text mining for the automatic identification and extraction of chemical entities. However, existing chemical NER systems primarily focus on scenarios with abundant training data, requiring significant human effort on annotations. This poses challenges for applications in the chemical field, such as catalysis, where many advancements have traditionally relied on trial-and-error investigations and incremental adjustment of variables. This hinders catalysis science and technology progress in addressing emerging energy and environmental crises. In this work, we propose a few-shot NER model that can quickly adapt to extract new types of chemical entities by using only a limited number of annotated examples. Our model employs a metric-learning approach to transfer entity similarity knowledge from high-resource chemical domains (with abundant annotations) to enable effective entity recognition in low-resource specialized domains (limited annotation). We validate the effectiveness of our model on a few-shot chemical NER benchmark built based on six existing chemical NER data sets. Experiments show that the proposed few-shot NER model can achieve reasonable performance with only 5 examples per entity type and shows consistent improvement as the number of examples increases. Furthermore, we demonstrate how the proposed model can be trained with large language model (LLM) annotated data, opening a new pathway for rapid adaptation of NER systems. Furthermore, our approach leverages the knowledge broadness of large language models for chemistry while distilling this knowledge into a lightweight model suitable for efficient and in-house use.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

A Model Based Approach to Extract Health Information from Textual Data

In current nuclear power plants (NPPs) a large amount of condition-based data is being generated and stored to assess and monitor component health and performance. The format of this data can be either numeric (e.g., pump vibration data) or textual (e.g., condition report which assess component health). While assessing component health from numeric data can be performed with a large variety of methods, the extraction of information from textual data still remains a challenge. Natural language processing (NLP) methods are starting to be deployed in current NPPs mainly to filter out incident reports (IRs) that are not safety related by employing supervised machine learning methods. However, these methods do not really provide the quantitative information that might be contained in IRs. This paper presents an approach to extract information from textual data (e.g., from IRs, maintenance reports) that is based on NLP data analytics methods coupled with model-based system engineer (MBSE) models. NLP methods are employed to perform syntactic and semantic analyses. Syntactic analysis analyzes the grammatical structure of a sentence; such analysis includes: part of speech (POS) tagging (i.e., identification of grammatic elements of each string - e.g., nouns, verbs), named entity recognition (i.e., identification of text entities - e.g., names, dates, events), and relation extraction (e.g., coreference resolution). On the other hand, semantic analysis is designed to analyze the logic structure of a sentence. Through a specific set of rules, our methods can identify whether a sentence contains health information of a component (e.g., degraded performance, anomaly behavior) or the causal relationship between two events (i.e., a cause-effect pair). An innovative element of our approach is that semantic analysis relies on MBSE models to identify links between textual elements. MBSE are diagrams designed to represent system and component dependencies (from both a form and functional point of view). In our approach, MBSE models emulate system engineer knowledge about component/system architecture. This paper presents in detail how the integration of NLP methods and MBSE models is performed. Few analysis examples focusing on centrifugal pumps are presented.

97 - MATHEMATICS AND COMPUTING

Machine learning and deep learning tools for the automated capture of cancer surveillance data

The National Cancer Institute and the Department of Energy strategic partnership applies advanced computing and predictive machine learning and deep learning models to automate the capture of information from unstructured clinical text for inclusion in cancer registries. Applications include extraction of key data elements from pathology reports, determination of whether a pathology or radiology report is related to cancer, extraction of relevant biomarker information, and identification of recurrence. With the growing complexity of cancer diagnosis and treatment, capturing essential information with purely manual methods is increasingly difficult. These new methods for applying advanced computational capabilities to automate data extraction represent an opportunity to close critical information gaps and create a nimble, flexible platform on which new information sources, such as genomics, can be added. This will ultimately provide a deeper understanding of the drivers of cancer and outcomes in the population and increase the timeliness of reporting. These advances will enable better understanding of how real-world patients are treated and the outcomes associated with those treatments in the context of our complex medical and social environment.

60 APPLIED LIFE SCIENCES

Understanding Event Trajectories Across Massive Temporal Datasets with Word Embeddings and Visualization

In collaboration with researchers from Virginia Tech, Savannah River National Laboratory has continued development of a natural language processing pipeline to identify and extract events of interest from massive open data sources in the domain of worldwide state-sponsored civil nuclear energy. The foundation of the pipeline is built on compass aligned temporal word embedding models, whereby contextual shifts are automatically identified by comparing keyword embedding vectors across successive time windows. Within the approach, a contextual shift indicates the occurrence of a potential event of interest. However, in such a broad topical domain that captures events at a global scale, across various life cycle stages, and across numerous different technology types, a user that is monitoring events may have broad interests in capturing many different event types with varying degrees of signal. As such, the quantity of information that may be returned from an automated event extraction pipeline can be substantial, requiring manual effort to sift through the information to identify any relevant bits of information. Therefore, a more streamlined workflow that aids in directing a user toward specific information at different points in time is necessary. The workflow presented here has been developed with this concept in mind, built on top of the initial prototype event extraction pipeline, whereby a user can analyze temporal text-based data sources at multiple different contextual levels to isolate key points in time and key subdomains captured within a data corpus. Using multiple corpuses that consist of approximately 7 million Tweets and 7 million news articles, the team has extended compass aligned temporal word embedding models to establish an interconnected and hierarchical structure that relates known key words of interest to documents, local topics (i.e., within a time window), and global topics across the corpuses. All of this information is packaged into a visual analytics system that is linked to the information extraction pipeline and enables a user to identify contextual information that describes the evolution of a high dimensional embedding space across time to isolate changes of interest and explore associated events. This report demonstrates the use of these analytics and a means to fuse information across multiple datasets.

97 MATHEMATICS AND COMPUTING

Ocpp 2.0.1. Interim Kpi Calculator

The project is split into four pieces. The first is a raw OCPP log parser. The second is a file splitter. The third is a message parser. The final piece is the Interim KPI calculator. The OCPP log parser was created from two different formats of raw OCPP 2.0.1 data. Its intended purpose is to extract device IDs and OCPP event messages from nontabular text logs. The parser looks for specific substrings in the logs to identify which of the two "standards" it should select from. The KPI generator does not perform any of its calculations in parallel. Instead, we opt for a naive batching approach. The splitter takes the file generated from the parser and creates many smaller files for each of the device IDs in the dataset. This allows the pandas queries in the log formatter to be iterate over a significantly smaller slice of data, increasing performance significantly. The message parser step takes messages from each of the files (containing distinct device IDs) and breaks the message out into pieces. The final result is a file with different columns specifying different attributes of the JSON message. The file is an aggregation of all different devices. This is the most complex portion of the code. The KPI calculator takes the parsed messages, as a single file, and calculates the KPI from that data. An excel file is produced with four sheets. These contain the metrics for Session Success, Charge Start Success, Charge End Success, and Charge Start Time. It includes the metrics for the different equations in the Interim KPI Implementation Guide as well as a weighted sum of the different equations for each KPI (excluding Charge End Success and Charge Start Time).

Quinn, Casey

Search for ${\text {Z}{}{}} {\text {Z}{}{}} $ and ${\text {Z}{}{}} {\text {H}{}{}} $ production in the ${\text {b}{}{}} {\bar{{\text {b}{}{}}}{}{}} {\text {b}{}{}} {\bar{{\text {b}{}{}}}{}{}} $ final state using proton-proton collisions at $\sqrt{s}=13\,\text {Te}\hspace{-.08em}\text {V} $

A search for ${\text {Z}{}{}} {\text {Z}{}{}} $ and ${\text {Z}{}{}} {\text {H}{}{}} $ production in the ${\text {b}{}{}} {\bar{{\text {b}{}{}}}{}{}} {\text {b}{}{}} {\bar{{\text {b}{}{}}}{}{}} $ final state is presented, where H is the standard model (SM) Higgs boson. The search uses an event sample of proton-proton collisions corresponding to an integrated luminosity of 133$\,\text {fb}^{-1}$ collected at a center-of-mass energy of 13$\,\text {Te}\hspace{-.08em}\text {V}$ with the CMS detector at the CERN LHC. The analysis introduces several novel techniques for deriving and validating a multi-dimensional background model based on control samples in data. A multiclass multivariate classifier customized for the ${\text {b}{}{}} {\bar{{\text {b}{}{}}}{}{}} {\text {b}{}{}} {\bar{{\text {b}{}{}}}{}{}} $ final state is developed to derive the background model and extract the signal. The data are found to be consistent, within uncertainties, with the SM predictions. The observed (expected) upper limits at 95% confidence level are found to be 3.8 (3.8) and 5.0 (2.9) times the SM prediction for the ${\text {Z}{}{}} {\text {Z}{}{}} $ and ${\text {Z}{}{}} {\text {H}{}{}} $ production cross sections, respectively.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

Scaling open-weight large language models for hydropower regulatory information extraction: A systematic analysis

Information extraction from regulatory and technical documents using large language models (LLMs) involves practical trade-offs between extraction quality and computational cost. We evaluate eight open-weight LLMs spanning 0.6B–70B parameters on hydropower licensing documents and report deployment-oriented evidence under a unified extraction schema and evaluation protocol. Across the model set, we observe clear scale-dependent trends in both baseline extraction quality and the effectiveness of reflective reasoning (self-checking) under our fixed-prompt, no-augmentation setting. Mid-scale models often provide a favorable balance of accuracy and efficiency, whereas the smallest models show limited or inconsistent gains from the reasoning variants tested. Larger models achieve the highest overall F1 scores but incur substantially greater compute and infrastructure requirements. We further find that reliability failure modes can distort conventional metrics in this domain: in particular, high recall can coincide with systematic extraction errors when models fabricate values for fields that are absent from the source text, underscoring the importance of conservative null handling and evidence-grounded evaluation. Overall, our study provides a reproducible resource–performance comparison for open-weight LLM-based extraction in hydropower regulatory documentation and offers practical guidance for model selection under different deployment constraints.

Evaluation protocol

Strangeness enhancement at its extremes: multiple (multi-)strange hadron production in pp collisions at \(\sqrt{s}=5.02\) TeV

The probability to observe a specific number of strange and multi-strange hadrons (nS), denoted as P(nS), is measured by ALICE at midrapidity (|y| < 0.5) in $$\sqrt{s}=5.02$$ TeV proton-proton (pp) collisions, dividing events into several multiplicity-density classes. Exploiting, for the first time, a technique based on counting the number of strange-particle candidates event-by-event, this measurement allows one to extend the study of strangeness production beyond the mean of the distribution. This constitutes a new test bench for production mechanisms, probing events with a large imbalance between strange and non-strange content. The analysis of a large-statistics data sample makes it possible to extract P(nS) up to a maximum nS of 7 for $${\text{K}}_{\text{S}}^{0}$$, 5 for Λ and $$\overline{\Lambda }$$, 4 for Ξ− and $${\overline{\Xi } }^{+}$$, and 2 for Ω− and $${\overline{\Omega } }^{+}$$. From this, the probability of producing strange hadron multiplets per event is calculated, thereby enabling the extension of the study of strangeness enhancement to extreme situations where several strange quarks hadronize in a single event at midrapidity. Moreover, comparing hadron combinations with different u and d quark compositions and equal overall s quark content, the contribution to the enhancement pattern coming from non-strangeness related mechanisms is isolated. The results are compared with state-of-the-art phenomenological models implemented in commonly used Monte Carlo event generators, including PYTHIA 8 Monash 2013, PYTHIA 8 with QCD-based Color Reconnection and Rope Hadronization (QCD-CR + Ropes), and EPOS LHC, which incorporates both partonic interactions and hydrodynamic evolution. These comparisons show that the new approach dramatically enhances the sensitivity to the different underlying physics mechanisms modeled by each generator.

Abualrob, I J

Optimizing Geospatial Assessments for Nuclear Safeguards Applications with Large Language Models

A multidisciplinary team at Argonne National Laboratory evaluated the ability of large language models (LLMs) to identify geographic locations from open-source text and assessed post-processing measures to strengthen the reliability of those extractions in support of international nuclear safeguards. The study focused on addressing challenges such as toponym ambiguity, imprecise descriptions, and misinformation, which often undermine the accuracy of LLM-derived geospatial assessments. By integrating authoritative geospatial datasets, employing rigorous validation techniques, and leveraging human-in-the-loop processes, the project aimed to enhance the precision, transparency, and reproducibility of geospatial localization workflows. The findings demonstrate that while LLMs exhibit significant potential for accelerating geospatial analysis, their outputs require systematic grounding and verification to ensure reliability in high-stakes applications. This work contributes to the broader field of geospatial intelligence and supports strategic objectives of international organizations such as the International Atomic Energy Agency (IAEA) and the U.S. Department of Energy (DOE).

97 MATHEMATICS AND COMPUTING

Production and Improved Separation of Therapeutic Radionuclides Tb-161, Er-165, and Lu-177

This project resulted in the synthesis and utility of new solid-phase extractants using diglycolamide (DGA) extractants grafted onto mesoporous silica. We used DGA ligands that offer multidentate coordination sites for lanthanides and offer pre-arranged binding sites that may facilitate radiolanthanide metal binding. Also, the Hunter graduate student worked at University of Utah on production and separation of 161 Tb with high specific activity resulting in a publication. These are reported in the full text upload.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA

ARCH: Large-scale knowledge graph via aggregated narrative codified health records analysis

Objective: Electronic health record (EHR) systems contain a wealth of clinical data stored as both codified data and free-text narrative notes (NLP). The complexity of EHR presents challenges in feature representation, information extraction, and uncertainty quantification. Here, to address these challenges, we proposed an efficient Aggregated naRrative Codified Health (ARCH) records analysis to generate a large-scale knowledge graph (KG) for a comprehensive set of EHR codified and narrative features. Methods: Using data from 12.5 million Veterans Affairs patients, ARCH first derives embedding vectors and generates similarities along with associated p-values to measure the strength of relatedness between clinical features with statistical certainty quantification. Next, ARCH performs a sparse embedding regression to remove indirect linkage between features to build a sparse KG. Finally, ARCH was validated on various clinical tasks, including detecting known relationships between entity pairs, predicting drug side effects, disease phenotyping, as well as sub-typing Alzheimer’s disease patients. Results: ARCH produces high-quality clinical embeddings and KG for over 60,000 codified and narrative EHR concepts. The KG and embeddings are visualized in the R-shiny powered web-API.3 ARCH achieved high accuracy in detecting EHR concept relationships, with AUCs of 0.926 (codified) and 0.861 (NLP) for similar EHR concepts, and 0.810 (codified) and 0.843 (NLP) for related pairs. It detected drug side effects with a 0.723 AUC, which improved to 0.826 after fine-tuning. Using both codified and NLP features, the detection power increased significantly. Compared to other methods, ARCH has superior accuracy and enhances weakly supervised phenotyping algorithms’ performance. Notably, it successfully categorized Alzheimer’s patients into two subgroups with varying mortality rates. Conclusion: The proposed ARCH algorithm generates large-scale high-quality semantic representations and knowledge graph for both codified and NLP EHR features, useful for a wide range of predictive modeling tasks.

Electronic health records

A Solid State Zwitterionic Plastic Crystal with High Static Dielectric Constant

The dielectric data in Figure 3, Figure 4, Figure S6 of the published paper was extracted from 2EOIMTSA-BDS-DATA .txt file. This file can be directly opened using a text file editor. It can also be imported to Excel/ Origin for further plotting and analysis. The G' and G'' in Figure 3 of the publihsed paper was plotted from data in file 2EOImTSA-temperature-sweep.xlsx. This file can be directly opend using Excel. The details of DFT simulations mentioned in Figure 2, Figure 7, and Figure S9 of the published paper are included in the DFT.zip file.

Huang, Zitan [Pennsylvania State University]

Fe(III) reducing bacterial activities in Old Woman Creek wetland sediments, June 2023

To evaluate the Fe(III) reducing microbiological activities in Old Woman Creek Nature Preserve (OWC) wetland sediments, we incubated OWC sediments under anoxic and oxic conditions and with or without Fe(III) amendment [as hydrous ferric oxide (HFO)]. No Fe(III) reduction was observed in heat-deactivated incubations. In non-sterile anoxic incubations, measurement of 0.5 M HCl-extractable Fe(II) indicated that Fe(III) reduction occurred in both Fe(III)-amended and -unamended incubations, indicating that abundant Fe(III) is associated with the OWC sediments. Little Fe(II) accumulated in solution, indicating that the most biogenic Fe(II) adsorbs to the sediments. When air was added to the headspace of non-sterile incubations, Fe(III) reduction was halted and any biogenic Fe(II) that accumulated was oxidized. These experiments were used to guide preparation and analyses of incubations to determine if electrochemical measuements can be used to detect microbiological activities in contrasting terminal electron accepting regimes (i.e., aerobic and Fe(III) reducing conditions). Data package includes methods and data from experiments, including dissolved anion concentrations, dissolved Fe(II) concentrations, and 0.5 M HCl-extractable Fe(II) concentrations. All files are either .txt or .csv and can be opened by any plain text editor application.

EARTH SCIENCE

Precision measurement of the B0 meson lifetime using B0→J/ψK∗0 decays with the ATLAS detector

A measurement of the B0$$B^{0}$$ meson lifetime using B0→J/ψK∗0$$ B^{0} \rightarrow J/\psi K^{*0} $$ decays in data from 13 TeV$$\text {TeV}$$ proton–proton collisions with an integrated luminosity of 140fb-1$$ 140~\mathrm {fb^{-1}} $$ recorded by the ATLAS detector at the LHC is presented. The measured effective lifetime is τ=1.5053±0.0012(stat.)±0.0035(syst.)ps.$$ \tau = 1.5053\pm 0.0012~\mathrm {(stat.)} \pm 0.0035~\mathrm {(syst.)~ps}. $$The average decay width extracted from the effective lifetime, using parameters from external sources, is Γd=0.6639±0.0005(stat.)±0.0016(syst.)±0.0038(ext.)ps-1,$$\begin{aligned} \Gamma _d = 0.6639\pm 0.0005~\mathrm {(stat.)} \pm 0.0016~\mathrm {(syst.)}\\ \pm 0.0038~\text {(ext.)} \text {~ps}^{-1}, \end{aligned}$$where the uncertainties are statistical, systematic and from external sources. The earlier ATLAS measurement of Γs$$\Gamma _s$$ in the Bs0→J/ψϕ$$B^{0}_{s} \rightarrow J/\psi \phi $$ decay was used to derive a value for the ratio of the average decay widths Γd$$\Gamma _d$$ and Γs$$\Gamma _s$$ for B0$$B^{0} $$ and Bs0$$B^{0}_{s} $$ mesons respectively, of ΓdΓs=0.9905±0.0022(stat.)±0.0036(syst.)±0.0057(ext.).$$ \frac{\Gamma _d }{\Gamma _s } = 0.9905\pm 0.0022~\text {(stat.)} \pm 0.0036~\text {(syst.)} \pm 0.0057~\text {(ext.)}. $$The measured lifetime, average decay width and decay width ratio are in agreement with theoretical predictions and with measurements by other experiments. This measurement provides the most precise result of the effective lifetime of the B0$$B^{0}$$ meson to date.

Aad, G

Measurement of the Drell–Yan forward-backward asymmetry and of the effective leptonic weak mixing angle in proton-proton collisions at $\sqrt{s}$ 13 TeV

The forward-backward asymmetry in Drell-Yan production and the effective leptonic electroweak mixing angle are measured in proton-proton collisions at $\sqrt{s}$ = 13 TeV, collected by the CMS experiment and corresponding to an integrated luminosity of 138 fb$^{-1}$. The measurement uses both dimuon and dielectron events, and is performed as a function of the dilepton mass and rapidity. The unfolded angular coefficient $A_4$ is also extracted, as a function of the dilepton mass and rapidity. Using the CT18Z set of parton distribution functions, we obtain $\sin^{2}\theta^\ell_\text{eff}$ = 0.23152 $\pm$ 0.00031, where the uncertainty includes the experimental and theoretical contributions. The measured value agrees with the standard model fit result to global experimental data. This is the most precise $\sin^{2}\theta^\ell_\text{eff}$ measurement at a hadron collider, with a precision comparable to the results obtained at LEP and SLD.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

Modeling and Simulation of Electrostatics of Ge$_{\text{1-x}}$Sn$_{\text{x}}$ Layers Grown on Ge Substrates

This work introduces a comprehensive simulation tool that provides a robust 1D Schrödinger – Poisson solver for modeling the electrostatics of heterostructures with an arbitrary number of layers, and non-uniform doping profiles along with the treatment of partial ionization of dopants at low temperatures. The effective masses are derived from the first-principles calculations. The solver is used to characterize three Ge 1-x Sn x /Ge heterostructures with non-uniform doping profiles and determine the subband structure at various temperatures. Here, the simulation results of the sheet carrier densities show excellent agreement with the experimentally extracted data, thus demonstrating the capabilities of the solver.

42 ENGINEERING

Explainable Synthesizability Prediction of Inorganic Crystal Polymorphs Using Large Language Models

Abstract We evaluate the ability of machine learning to predict whether a hypothetical crystal structure can be synthesized and explain those predictions to scientists. Fine‐tuned large language models (LLMs) trained on a human‐readable text description of the target crystal structure perform comparably to previous bespoke convolutional graph neural network methods, but better prediction quality can be achieved by training a positive‐unlabeled learning model on a text‐embedding representation of the structure. An LLM‐based workflow can then be used to generate human‐readable explanations for the types of factors governing synthesizability, extract the underlying physical rules, and assess the veracity of those rules. These explanations can guide chemists in modifying or optimizing non‐synthesizable hypothetical structures to make them more feasible for materials design.

Kim, Seongmin [Department of Chemical and Biologic