Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Unstructured”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Towards the holistic design of alloys with large language models

Large language models are very effective at solving general tasks, but can also be useful in materials design and extracting and using information from the scientific literature and unstructured corpora. For instance, in the domain of alloy design and manufacturing, they can expedite the materials design process and enable the inclusion of holistic criteria.

36 MATERIALS SCIENCE↗

Decoding substance use disorder severity from clinical notes using a large language model

Substance use disorder (SUD) poses a major concern due to its detrimental effects on health and society. SUD identification and treatment depend on a variety of factors such as severity, co-determinants (e.g., withdrawal symptoms), and social determinants of health. Existing diagnostic coding systems used by insurance providers, like the International Classification of Diseases (ICD-10), lack granularity for certain diagnoses, but American clinicians will add this granularity (as that found within the Diagnostic and Statistical Manual of Mental Disorders classification or DSM-5) as supplemental unstructured text in clinical notes. Traditional natural language processing (NLP) methods face limitations in accurately parsing such diverse clinical language. Large language models (LLMs) offer promise in overcoming these challenges by adapting to diverse language patterns. This study investigates the application of LLMs for extracting severity-related information for various SUD diagnoses from clinical notes. We propose a workflow employing zero-shot learning of LLMs with carefully crafted prompts and post-processing techniques. Through experimentation with Flan-T5, an open-source LLM, we demonstrate its superior recall compared to the rule-based approach. Focusing on 11 categories of SUD diagnoses, we show the effectiveness of LLMs in extracting severity information, contributing to improved risk assessment and treatment planning for SUD patients.

60 APPLIED LIFE SCIENCES↗

Advanced silicone materials for soft actuator applications

Soft actuators are at the forefront of the innovation tide in medicine, manufacturing, and aerospace because they are able to mimic the behavior of biological tissue and adapt to complex, unstructured environments. Of all the materials used, silicone-based elastomers have drawn enormous attention since they offer a superb combination of mechanical flexibility, biocompatibility, thermal stability, and long-term durability. In the past few years, there has also been a rapid pace of material evolution, additive manufacturing, and biointegration that has enhanced the performance and applications of silicone-based soft actuators. However, there is no focused and timely review compiling these advances. This review seeks to address that need by critically discussing recent advancements in advanced silicone materials, exploring new fabrication methodologies, and discussing emerging applications that range from wearable devices to implantable robotics. We also present suggestions for directions and the problems which must be addressed in order to further develop the performance and potential of silicone-based soft actuators, justifying the relevance and urgency of this effort.

biocompatibility↗

MAPS: the MFEM Anisotropic Plasma Solver

Simulating magnetically confined fusion plasmas presents a uniquely challenging problem due to the nonlinear anisotropic heat conduction. We introduce the MAPS (MFEM Anisotropic Plasma Solver) tool, which uses a high-order finite element method to compute transport solutions on unstructured meshes. We show results for a set of three 2-D verification tests, two of which demonstrate the expected convergence properties for various mesh resolutions and polynomial degrees. We then discuss the convergence rate for the third test.

Barnett, Rhea [ORNL] (ORCID:0000000317527979)↗

Dynamic in-context learning with conversational models for data extraction and materials property prediction

The advent of natural language processing and large language models (LLMs) has revolutionized the extraction of data from unstructured scholarly papers. However, ensuring data trustworthiness remains a significant challenge. In this paper, we introduce PropertyExtractor, an open-source tool that leverages advanced conversational LLMs such as Google gemini-pro and OpenAI gpt-4, blends zero-shot with few-shot in-context learning, and employs engineered prompts for the dynamic refinement of structured information hierarchies—enabling autonomous, efficient, scalable, and accurate identification, extraction, and verification of material property data. Our tests on material data demonstrate precision and recall that exceed 95% with an error rate of ∼9%, highlighting the effectiveness and versatility of the toolkit. Finally, databases for 2D material thicknesses, a critical parameter for device integration, and energy bandgap values are developed using PropertyExtractor. In particular, for the thickness database, the rapid evolution of the field has outpaced both experimental measurements and computational methods, creating a significant data gap. Our work addresses this gap and showcases the potential of PropertyExtractor as a reliable and efficient tool for the autonomous generation of various material property databases, advancing the field.

Ekuma, Chinedu E. (ORCID:0000000258527556)↗

Soft x-ray high-harmonic generation in an anti-resonant hollow core fiber driven by a 3 μ m ultrafast laser

High-harmonic upconversion driven by a mid-infrared femtosecond laser can generate coherent soft x-ray beams in a tabletop-scale setup. Here, we report on a compact ytterbium-pumped optical parametric chirped pulse amplifier (OPCPA) laser system seeded by an all-fiber front-end and employing periodically poled lithium niobate (PPLN) nonlinear media operated near the pulse fluence limits of current commercially available PPLN crystals. The OPCPA delivers 3 µm wavelength pulses with 775 µJ energy at 1 kHz repetition rate, with transform-limited 120 fs pulse duration, diffraction-limited beam quality, and ultrahigh 0.33% rms energy stability over >18 h. Using this laser, we generate soft x-ray high harmonics (HHG) in argon gas by focusing into a low-loss, high-pressure gas-filled anti-resonant hollow core fiber (ARHCF), generating coherent light at photon energies up to the argon L-edge (250 eV) and carbon K-edge (284 eV), with high beam quality and ∼1% rms energy stability. This work demonstrates soft x-ray HHG in a high-efficiency guided-wave phase matched geometry, overcoming the high losses inherent to mid-IR propagation in unstructured waveguides, or the short interaction lengths of gas cells or jets. The ARHCF can operate in the long term without damage and with the repetition rate, stability, and robustness required for demanding applications in spectromicroscopy and imaging. Finally, we discuss routes for further optimizing the soft x-ray HHG flux by driving He at higher laser intensities using either the signal (1.5 μm) or idler wavelengths (3 μm).

Femtosecond lasers↗

A Second Moment Method for k -Eigenvalue Acceleration with Continuous Diffusion and Discontinuous Transport Discretizations

The second moment method is a linear acceleration technique that couples the transport equation to a diffusion equation with transport-dependent additive closures. The resulting low-order diffusion equation can be discretized independent of the transport discretization, unlike diffusion synthetic acceleration, and is symmetric positive definite, unlike quasidiffusion. While this method has been shown to be comparable to quasidiffusion in iterative performance for fixed source and time-dependent problems, it is largely unexplored as an eigenvalue problem acceleration scheme due to the belief that the resulting inhomogeneous source makes the problem ill posed. Recently, a preliminary feasibility study was performed on the second moment method for eigenvalue problems. The results suggested comparable performance to quasidiffusion and more robust performance than diffusion synthetic acceleration. This work extends the initial study to more realistic reactor problems using state-of-the-art discretization techniques. Finally, the results in this paper show that the second moment method is more computationally efficient than its alternatives on complex reactor problems with unstructured meshes.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

One-shot omnidirectional pressure integration through matrix inversion

In this work, we present a method to perform 2D and 3D omnidirectional pressure integration from velocity measurements with a single-iteration matrix inversion approach. This work builds upon our previous work, where the rotating parallel ray approach was extended to the limit of infinite rays by taking continuous projection integrals of the ray paths and recasting the problem as an iterative matrix inversion problem. This iterative matrix equation is now 'fast-forwarded' to the 'infinity' iteration, leading to a different matrix equation that can be solved in a single step, thereby presenting the same computational complexity as the Poisson equation. We observe computational speedups of ~10 6 when compared to brute-force omnidirectional integration methods, enabling the treatment of grids of ~10 9 points and potentially even larger in a desktop setup at the time of publication. Further examination of the boundary conditions of our one-shot method shows that omnidirectional pressure integration implements a boundary condition where the boundary points are treated as interior points to the extent that information is available. Finally, we show how the method can be extended from the regular grids typical of particle image velocimetry to the unstructured meshes characteristic of particle tracking velocimetry data.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Response of salt intrusion in a tidal estuary to regional climatic forcing

Abstract Salinity distribution in a large tidal estuary is subject to estuarine adjustment under the influences of multiple physical drivers such as freshwater pulses and sea level rise, and is crucial to upstream water quality, aquaculture, and ecosystem functions of the estuary. To better understand the estuarine salinity response to climate change, the unstructured-grid Finite Volume Community Ocean Model was implemented to simulate the salt intrusion in the Delaware Bay Estuary. The model was first validated by multiple observational data sets and subsequently applied in an idealized setting to examine the response of salt front to freshwater pulses in high flow conditions, followed by a long-term drought condition supported by a multi-decadal streamflow drought analysis in the estuary. The model results showed that after the freshwater pulses the salt front location moved further upstream with sea level rise. Under the simulated long-term drought condition, the adjustment timescale of salt intrusion varies nonlinearly with sea level rise. With a significant increase in sea level rise, the adjustment timescale starts to decrease. This shift suggests a transition into a different regime where the estuary becomes more stratified, as indicated by an increasing bulk Simpson number with rising sea levels.

54 ENVIRONMENTAL SCIENCES↗

32 examples of LLM applications in materials science and chemistry: towards automation, assistants, agents, and accelerated scientific discovery

Abstract Large language models (LLMs) are reshaping many aspects of materials science and chemistry research, enabling advances in molecular property prediction, materials design, scientific automation, knowledge extraction, and more. Recent developments demonstrate that the latest class of models are able to integrate structured and unstructured data, assist in hypothesis generation, and streamline research workflows. To explore the frontier of LLM capabilities across the research lifecycle, we review applications of LLMs through 32 total projects developed during the second annual LLM hackathon for applications in materials science and chemistry, a global hybrid event. These projects spanned seven key research areas: (1) molecular and material property prediction, (2) molecular and material design, (3) automation and novel interfaces, (4) scientific communication and education, (5) research data management and automation, (6) hypothesis generation and evaluation, and (7) knowledge extraction and reasoning from the scientific literature. Collectively, these applications illustrate how LLMs serve as versatile predictive models, platforms for rapid prototyping of domain-specific tools, and much more. In particular, improvements in both open source and proprietary LLM performance through the addition of reasoning, additional training data, and new techniques have expanded effectiveness, particularly in low-data environments and interdisciplinary research. As LLMs continue to improve, their integration into scientific workflows presents both new opportunities and new challenges, requiring ongoing exploration, continued refinement, and further research to address reliability, interpretability, and reproducibility.

Computer Science↗

SetBERT: the deep learning platform for contextualized embeddings and explainable predictions from high-throughput sequencing

MOTIVATION: High-throughput sequencing (HTS) is a modern sequencing technology used to profile microbiomes by sequencing thousands of short genomic fragments from the microorganisms within a given sample. This technology presents a unique opportunity for artificial intelligence to comprehend the underlying functional relationships of microbial communities. However, due to the unstructured nature of HTS data, nearly all computational models are limited to processing DNA sequences individually. This limitation causes them to miss out on key interactions between microorganisms, significantly hindering our understanding of how these interactions influence the microbial communities as a whole. Furthermore, most computational methods rely on post-processing of samples which could inadvertently introduce unintentional protocol-specific bias. RESULTS: Addressing these concerns, we present SetBERT, a robust pre-training methodology for creating generalized deep learning models for processing HTS data to produce contextualized embeddings and be fine-tuned for downstream tasks with explainable predictions. By leveraging sequence interactions, we show that SetBERT significantly outperforms other models in taxonomic classification with genus-level classification accuracy of 95%. Furthermore, we demonstrate that SetBERT is able to accurately explain its predictions autonomously by confirming the biological-relevance of taxa identified by the model. AVAILABILITY AND IMPLEMENTATION: All source code is available at https://github.com/DLii-Research/setbert. SetBERT may be used through the q2-deepdna QIIME 2 plugin whose source code is available at https://github.com/DLii-Research/q2-deepdna.

Ludwig, David W↗

Scalable edge clustering of dynamic graphs via weighted line graphs

Timestamped relational datasets consisting of records (or connections) between pairs of entities are ubiquitous in network science. For applications like peer-to-peer communication, email, various social network interactions, and computer network security, it is useful to organize these records into groups based on how and when they are occurring. Weighted line graphs offer a natural way to model how records are related in such datasets but for large real-world graph topologies, building and utilizing the line graph is prohibitively expensive. Here, we present the framework to cluster the edges of a dynamic graph via the associated line graph that contains two major contributions. The first is a method to work with the line graph implicitly and the second is a distributed scale implementation of an agglomerative hierarchical graph clustering algorithm. We outline a novel hierarchical dynamic graph edge clustering approach that efficiently breaks massive relational datasets into small sets of edges containing events at various timescales. This is in stark contrast to traditional graph clustering algorithms that prioritize highly connected (clique-like) community structures. Our approach relies on constructing a sufficient subgraph of a weighted line graph and applying a hierarchical agglomerative clustering. This approach is related to scalable techniques from spatial clustering, nonlinear-dimension reduction, topological data analysis, and draws particular inspiration from HDBSCAN. As an edge clustering, this method yields an overlapping node clustering. Our algorithm is parallelizable and we demonstrate efficient clustering of a billion-scale, real-world dynamic graph into small edge sets that correlate in topology and time. The entire clustering process for a graph with tens of billions of edges takes just a few minutes of run time on 256 nodes of a distributed compute environment. We argue how the output of the edge clustering is useful for a multitude of data visualization and powerful machine learning tasks, both involving the original massive dynamic graph data and metadata associated with the nodes and edges. Finally, we describe how this approach can be extended to dynamic hypergraphs and dynamic graphs/hypergraphs with unstructured data living on vertices and edges.

Data Analysis↗

Machine learning and deep learning tools for the automated capture of cancer surveillance data

The National Cancer Institute and the Department of Energy strategic partnership applies advanced computing and predictive machine learning and deep learning models to automate the capture of information from unstructured clinical text for inclusion in cancer registries. Applications include extraction of key data elements from pathology reports, determination of whether a pathology or radiology report is related to cancer, extraction of relevant biomarker information, and identification of recurrence. With the growing complexity of cancer diagnosis and treatment, capturing essential information with purely manual methods is increasingly difficult. These new methods for applying advanced computational capabilities to automate data extraction represent an opportunity to close critical information gaps and create a nimble, flexible platform on which new information sources, such as genomics, can be added. This will ultimately provide a deeper understanding of the drivers of cancer and outcomes in the population and increase the timeliness of reporting. These advances will enable better understanding of how real-world patients are treated and the outcomes associated with those treatments in the context of our complex medical and social environment.

60 APPLIED LIFE SCIENCES↗

TomoPyUI : a user-friendly tool for rapid tomography alignment and reconstruction

The management and processing of synchrotron and neutron computed tomography data can be a complex, labor-intensive and unstructured process. Users devote substantial time to both manually processing their data ( i.e. organizing data/metadata, applying image filters etc. ) and waiting for the computation of iterative alignment and reconstruction algorithms to finish. In this work, we present a solution to these problems: TomoPyUI , a user interface for the well known tomography data processing package TomoPy . This highly visual Python software package guides the user through the tomography processing pipeline from data import, preprocessing, alignment and finally to 3D volume reconstruction. The TomoPyUI systematic intermediate data and metadata storage system improves organization, and the inspection and manipulation tools (built within the application) help to avoid interrupted workflows. Notably, TomoPyUI operates entirely within a Jupyter environment. Herein, we provide a summary of these key features of TomoPyUI , along with an overview of the tomography processing pipeline, a discussion of the landscape of existing tomography processing software and the purpose of TomoPyUI , and a demonstration of its capabilities for real tomography data collected at SSRL beamline 6-2c.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

GraphAide: Advanced Graph-Assisted Query and Reasoning System

Curating knowledge from multiple siloed sources that contain both structured and unstructured data is a major challenge in many real-world applications. Pattern matching and querying represent fundamental tasks in modern data analytics that leverage this curated knowledge. The development of such applications necessitates overcoming several research challenges, including data extraction, named entity recognition, data modeling, and designing query interfaces. Moreover, the explainability of these functionalities is critical for their broader adoption. The emergence of Large Language Models (LLMs) has accelerated the development lifecycle of new capabilities. Nonetheless, there is an ongoing need for domain-specific tools tailored to user activities. The creation of digital assistants has gained considerable traction in recent years, with LLMs offering a promising avenue to develop such assistants utilizing domain-specific knowledge and assumptions. In this context, we introduce an advanced query and reasoning system, GraphAide, which constructs a knowledge graph (KG) from diverse sources and allows to query and reason over the resulting KG. GraphAide harnesses both the KG and LLMs to rapidly develop domain-specific digital assistants. It integrates design patterns from retrieval augmented generation (RAG) and the semantic web to create an agentic LLM application. GraphAide underscores the potential for streamlined and efficient development of specialized digital assistants, thereby enhancing their applicability across various domains.

Purohit, Sumit [BATTELLE (PACIFIC NW LAB)] (ORCID:↗

Investigating Resilience of Loops in HPC Programs: A Semantic Approach with LLMs

Soft errors have become one of the major concerns for the error resilience of the HPC applications as those errors may cause HPC applications to generate serious outcomes such as silent data corruptions (SDCs). Protecting the applications from soft errors is an essential while challenging task. Among different approaches, obtaining a profound understanding of the resilience proneness of an application is very important to devise efficient error detection and recovery strategies. Given the scale of the HPC applications both in the code size and execution time, there are often cases that the error propagation analysis on such applications would produce a massive volume of unstructured data, which requires a significant amount of efforts, to process and to obtain indicating actions towards error protection. In this paper, we present a control-flow based visual analysis framework to help the users conduct error propagation analysis and identify the critical sections of a program that may have a higher likelihood of leading to erroneous outcomes when affected by the control flow related errors. We also design and implement the scalable visualization framework - ResilienceVis that efficiently and effectively visualizes the affected program states under errors and the propagation traces for an application in a user-friendly manner, and eventually, we combine the analysis and visualization to exhibit the error-proneness of the different sections of applications.

Jiang, Hailong↗

Towards Automatically Matching Security Advisories to CPEs: String Similarity-based Vendor Matching

When a vulnerability is reported by the National Vulnerability Database (NVD), affected products are listed in the structured Common Platform Enumeration (CPE) format. Unfortunately, if the vulnerability is in a software library (e.g., Log4j), it will not include CPEs for each product containing that library. In these cases, security operators need to manually read the vendor's or third-party security advisories to see if their product is affected. However, these advisories do not report affected products in a structured format, which prevents automated processing, This paper makes the first effort towards automatically constructing structured CPEs for the vulnerable products in a non-NVD security advisory from the unstructured data in the advisory. Since this is a very challenging problem, this paper specifically focuses on the initial but key step of matching the un-structured vendor names in security advisories to the structured vendor representations in the standard CPE format. We explore the feasibility of using string similarity to solve the problem. The basic idea is to compare a vendor name from the non-NVD advisory with each vendor in the official CPE dictionary. The CPE vendor with the highest similarity score to the advisory's vendor will be considered as the match. We first conduct an experimental, comparative study of multiple mainstream string similarity metrics for this matching problem. To improve the performance, we then design a new string similarity metric that is adapted from an existing metric by weighing different tokens in the advisory's vendor name differently.

McClanahan, Kylie↗

A Perspective on Data and Privacy for AI in Healthcare [Industrial and Governmental Activities]

As large language models continue to push the bounds of AI model size, they are also being trained on unprecedented volumes of data. While individual hospitals are estimated to produce petabytes of data per year, only a small fraction is currently being used for developing AI models. Additionally, with such data resources available, healthcare is well-positioned to benefit from the current trends in AI. Moreover, the inherently multi-modal and longitudinal nature of clinical data – from omics to imaging to unstructured notes – provides a fertile ground for the development and application of cutting-edge architectures like foundation models.

Gounley, John [Oak Ridge National Laboratory (ORNL↗