Engineering PapersSearch

SEARCH · Engineering Papers

Results for “data extraction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Dynamic in-context learning with conversational models for data extraction and materials property prediction

The advent of natural language processing and large language models (LLMs) has revolutionized the extraction of data from unstructured scholarly papers. However, ensuring data trustworthiness remains a significant challenge. In this paper, we introduce PropertyExtractor, an open-source tool that leverages advanced conversational LLMs such as Google gemini-pro and OpenAI gpt-4, blends zero-shot with few-shot in-context learning, and employs engineered prompts for the dynamic refinement of structured information hierarchies—enabling autonomous, efficient, scalable, and accurate identification, extraction, and verification of material property data. Our tests on material data demonstrate precision and recall that exceed 95% with an error rate of ∼9%, highlighting the effectiveness and versatility of the toolkit. Finally, databases for 2D material thicknesses, a critical parameter for device integration, and energy bandgap values are developed using PropertyExtractor. In particular, for the thickness database, the rapid evolution of the field has outpaced both experimental measurements and computational methods, creating a significant data gap. Our work addresses this gap and showcases the potential of PropertyExtractor as a reliable and efficient tool for the autonomous generation of various material property databases, advancing the field.

Ekuma, Chinedu E. (ORCID:0000000258527556)

Transition Metals Separation with Commercial Neutral Extractants – A Review

The increasing use of extraction chromatography resins across fields such as hydrometallurgy, nuclear medicine, and environmental analysis has created a need for a deeper understanding of their interactions with transition metals. Despite extensive research on f-element separations, the behavior of transition metals in these systems remains relatively understudied. This review provides a comprehensive overview of the current state of knowledge on the extraction behavior of transition metals with neutral extractants, including TODGA, TEHDGA, TBP, and CMPO, and their corresponding resins, such as DGA, BDGA, UTEVA, TBP, and TRU. The review summarizes extraction data, extracted complex coordination environments, separation reaction stoichiometries, and associated thermodynamics, highlighting inconsistencies and knowledge gaps in the literature. The study emphasizes the need for further research using spectroscopy and computational methods to elucidate extraction mechanisms and to improve the efficiency and selectivity of transition metal separations. By identifying areas for future research and development, this review aims to stimulate advancements in the field and promote the development of innovative separation technologies. The implications of this research are far-reaching, with potential applications in nuclear waste management, nuclear forensics, metal recovery, and environmental remediation. Overall, this review provides a foundation for future studies on the extraction of transition metals using neutral extractants and resins.

Wall, Nathalie A.

Leveraging Large Language Models for Real-World Data Evidence: A Framework for Automated Treatment Extraction and Data Harmonization

Background: The ability to comprehensively collect treatment information from cancer patient medical records would enable studies to evaluate real-world benefits and risks tied to specific treatments. Currently, it is difficult to system- atically collect high-quality treatment information because it is often stored in unstructured text. Manually extracting and standardizing drug and regimen data is time-intensive. Recent advances in large language models (LLMs) offer a potential solution for automated extraction of structured treatment information from clinical text. Objective: This study systematically evaluates the utility of four LLMs from the Llama family for automated extraction of oncology treatment information from clinical text. This information can guide researchers using cancer registry data to provide insights into cancer care and outcomes beyond clinical trials. Methods: Four instruction-tuned Llama models with varying parameter counts (1B, 3B, 8B, and 70B) were evaluated for their ability to extract treatment information from clinical documents. A unified oncology knowledge base integrating seven major public data sources was developed to standardize and normalize extracted entities—a critical step for harmonizing data from diverse sources. Extracted treatment data were compared against expert-annotated ground truth. Model performance was assessed using accuracy metrics (Precision, Recall, F1-Score) and opera- tional feasibility metrics, including processing speed and structural compliance of the output. Results: A strong positive correlation was observed between model size and extraction accuracy. F1-score improved from 0.609 for the 1B model to 0.710 (3B), 0.807 (8B), and 0.828 (70B). While larger models demonstrated superior accuracy and compliance, they incurred higher computational costs. The modest performance difference between 8B and 70B suggests diminishing returns with increasing model size. Conclusions: LLMs represent a viable technology for automating oncology treatment extraction. The 8B-parameter model emerged as a highly effective option, balancing high accuracy and computational efficiency. Selecting an appropriate LLM for deployment in cancer registries involves a trade-off between desired accuracy and available operational resources. Harmonizing extracted entities with the oncology knowledge base facilitates standardized integration into common data models, enhancing data quality for real-world evidence analyses.

artificial intelligence

TropiRoot 1.0: Database of tropical root characteristics across environments

Tropical ecosystems contain the world's largest biodiversity of vascular plants. Yet, our understanding of tropical functional diversity and its contribution to global diversity patterns is constrained by data availability. This discrepancy underscores an urgent need to bridge data gaps by incorporating comprehensive tropical root data into global datasets. Here, we provide a database of tropical root characteristics. This new database, TropiRoot 1.0, will be instrumental in evaluating an array of hypotheses pertaining to root functional ecology and plant biogeography, both within the tropics and relative to other global biomes. The data compilation was conducted by the TropiRoot Initiative, in partnership with the Fine-Root Ecology Database (FRED) and the Global Root Trait (GRooT) database, Colorado State University (CSU) and the Smithsonian Tropical Research Institute (STRI). Literature search and data extraction were conducted between 2020 and 2024. Literature was identified using Web of Science, Scopus, and complemented using the expert knowledge of members of TropiRoot. To provide broad environmental and geographical distributions, literature searches included root characteristics (traits) across global change drivers, natural gradients, and from different continents. We adopted FRED standardized data columns and streamlined the format to enhance accessibility for data extraction across various user groups. This optimized framework resulted in a smaller, yet comprehensive datasheet. To make the database compatible with other global root trait initiatives, column identification was standardized following the codes provided by FRED. These efforts culminated in data extracted from 104 new sources, resulting in more than 8000 rows of data (either species or community data). Most of the data in TropiRoot 1.0 include root characteristics such as root biomass, morphology, root dynamics, mass fraction, architecture, anatomy, physiology, and root chemistry. This initiative represents a 30% increase in the currently available data for tropical roots in FRED. TropiRoot 1.0 contains root characteristics from 25 different countries, where seven are located in Asia, six in South America, five in Central America and the Caribbean, four in Africa, two in North America, and 1 in Oceania. Due to the volume of data, when ancillary data were available, including soil data, these data were either extracted and included in the database or its availability was recorded in an additional column. Multiple contributors checked the entries for outliers during the collation process to ensure data quality. For text-based observations, we examined all cells to ensure that their content relates to their specific categories. For numerical observations, we ordered each numerical value from least to greatest and plotted the values, checking apparent outliers against the data in their respective sources and correcting or removing incorrect or impossible values. Some data (soil and aboveground) have different columns for the same variable presented in different units, including originally published units, but root characteristics data had units converted to match those reported in FRED. By filling a gap from global databases, TropiRoot 1.0 expands our knowledge of otherwise so far underrepresented regions and our ability to assess global trends. This advancement can be used to improve tropical forest representation in vegetation models. The data are freely available and should be cited when used.

FRED

A Database of Stress-Strain Properties Auto-generated from the Scientific Literature using ChemDataExtractor

Abstract There has been an ongoing need for information-rich databases in the mechanical-engineering domain to aid in data-driven materials science. To address the lack of suitable property databases, this study employs the latest version of the chemistry-aware natural-language-processing (NLP) toolkit, ChemDataExtractor, to automatically curate a comprehensive materials database of key stress-strain properties. The database contains information about materials and their cognate properties: ultimate tensile strength, yield strength, fracture strength, Young’s modulus, and ductility values. 720,308 data records were extracted from the scientific literature and organized into machine-readable databases formats. The extracted data have an overall precision, recall and F-score of 82.03%, 92.13% and 86.79%, respectively. The resulting database has been made publicly available, aiming to facilitate data-driven research and accelerate advancements within the mechanical-engineering domain.

Kumar, Pankaj

Computational epidemiological tools for pandemic analysis, understanding, and response

This suite of software tools is being developed to enhance and analyze computational epidemiological models that incorporate realistic disease dynamics and human behavior, with the goal of supporting epidemic and pandemic response. Specifically, the tools enable data analysis, feature extraction, data synthesis, machine learning model development, and prediction of key public health outcomes, such as cases, hospitalizations, deaths, and behavioral responses, for airborne infectious diseases like COVID-19 and influenza.

Butts, David

Physical properties, internal structure, and the three‐dimensional petrography of CI chondrites

physical properties and the nature of their breccation, we investigated nine samples of the Ivuna and Orgueil CI chondrites ranging in size from 1 mm to 4 cm in approximate diameter. The combined mass of unique material investigated in this work is 113 g. For our investigations, we use ideal gas pycnometry, 3-D laser scanning, x-ray computed microtomography (μCT), and accompanying digital data extraction techniques. We found that the bulk density of the samples ranged from 1.61 to 2.10 g cm −3 . Larger samples tend to have a lower bulk density. Grain density (ranging from 2.44 to 2.55 g cm −3 ) is significantly less variable than the bulk density in our samples and the quantity of porosity (ranging from 14.6% to 33.8%) is the dominant factor in determining the bulk density of CI chondrite material. Our μCT results show that the visible porosity across all sizes of our CI chondrite samples is in the form of cracks, but these cracks can account for less than two-thirds of the porosity in the CI chondrites. Other porosity is not visible, even at μCT resolutions of 2.7 μm voxel edge −1 and we conclude that it is sub-micron in nature. It is not clear if the cracks seen in our samples are indigenous to the chondrites or are a result of terrestrial processes. We also find that the CI chondrites are excellent examples of the fractal-like nature of brecciation, where clasts can be observed at all scales we imaged. The breccias are composed of sub-equant-shaped and sub-rounded-textured clasts like melt-free impact breccias on other solar system bodies. From our μCT volume and digital data extraction, we determine that the Ivuna CI chondrite breccia is organized: the mostly sub-equant clasts within our ~2 cm chunk of Ivuna have a mean diameter of 1.33 mm and their aligned longest axes define a lineation structure. We speculate that the lineation was imparted after fragmentation of the clasts by slight shear on the parent asteroid which could be the result of seismic-related granular flow or mild non-axial impact-related compaction. These data will help to place returned asteroidal material from asteroids 162173 Ryugu and 101955 Bennu and the CI chondrites into a mutual geological context.

CI chondrite

GraphAide: Advanced Graph-Assisted Query and Reasoning System

Curating knowledge from multiple siloed sources that contain both structured and unstructured data is a major challenge in many real-world applications. Pattern matching and querying represent fundamental tasks in modern data analytics that leverage this curated knowledge. The development of such applications necessitates overcoming several research challenges, including data extraction, named entity recognition, data modeling, and designing query interfaces. Moreover, the explainability of these functionalities is critical for their broader adoption. The emergence of Large Language Models (LLMs) has accelerated the development lifecycle of new capabilities. Nonetheless, there is an ongoing need for domain-specific tools tailored to user activities. The creation of digital assistants has gained considerable traction in recent years, with LLMs offering a promising avenue to develop such assistants utilizing domain-specific knowledge and assumptions. In this context, we introduce an advanced query and reasoning system, GraphAide, which constructs a knowledge graph (KG) from diverse sources and allows to query and reason over the resulting KG. GraphAide harnesses both the KG and LLMs to rapidly develop domain-specific digital assistants. It integrates design patterns from retrieval augmented generation (RAG) and the semantic web to create an agentic LLM application. GraphAide underscores the potential for streamlined and efficient development of specialized digital assistants, thereby enhancing their applicability across various domains.

Purohit, Sumit [BATTELLE (PACIFIC NW LAB)] (ORCID:

MaTableGPT: GPT‐Based Table Data Extractor from Materials Science Literature

Abstract Efficiently extracting data from tables in the scientific literature is pivotal for building large‐scale databases. However, the tables reported in materials science papers exist in highly diverse forms; thus, rule‐based extractions are an ineffective approach. To overcome this challenge, the study presents MaTableGPT, which is a GPT‐based table data extractor from the materials science literature. MaTableGPT features key strategies of table data representation and table splitting for better GPT comprehension and filtering hallucinated information through follow‐up questions. When applied to a vast volume of water splitting catalysis literature, MaTableGPT achieves an extraction accuracy (total F1 score) of up to 96.8%. Through comprehensive evaluations of the GPT usage cost, labeling cost, and extraction accuracy for the learning methods of zero‐shot, few‐shot, and fine‐tuning, the study presents a Pareto‐front mapping where the few‐shot learning method is found to be the most balanced solution owing to both its high extraction accuracy (total F1 score >95%) and low cost (GPT usage cost of 5.97 US dollars and labeling cost of 10 I/O paired examples). The statistical analyses conducted on the database generated by MaTableGPT revealed valuable insights into the distribution of the overpotential and elemental utilization across the reported catalysts in the water splitting literature.

Yi, Gyeong Hoon [Computational Science Research Ce

Data from TropiRoot 1.0 database: tropical root characteristics across environments

TropiRoot 1.0 is a new tropical root database with root characteristics across environment gradients. It has data extracted from 104 new sources, resulting in more than 8000 rows of data (either species or community data). Most of the data in TropiRoot 1.0 includes root characteristics such as root biomass, morphology, root dynamics, mass fraction, architecture, anatomy, physiology and root chemistry. This initiative represents an approximately 30% increase in the currently available data for tropical roots in the Fine Root Ecology Database (FRED). TropiRoot 1.0, contains root characteristics from 25 different countries where seven are located in Asia, six in South America, five in Central America and the Caribbean, four in Africa, two in North America, and 1 in Oceania. Due to the volume of data, when ancillary data was available, including soil data, these data was either extracted and included in the database or their availability was recorded in an additional column. Multiple contributors checked the entries for outliers during the collation process to ensure data quality. For text-based observations, we examined all cells to ensure that their content relates to their specific categories. For numerical observations, we ordered each numerical value from least to greatest and plotted the values, checking apparent outliers against the data in their respective sources and correcting or removing incorrect or impossible values. Some data (soil and aboveground) have different columns for the same variable presented in different units, including originally published units, but root characteristics data had units converted to match the ones reported in FRED. By filling a gap from global databases, TropiRoot 1.0 expands our knowledge of otherwise so far underrepresented regions, and our ability to assess global trends. This advancement can be used to improve tropical forest representation in vegetation models.

54 ENVIRONMENTAL SCIENCES

Emergent Nanostructure and Ion Transport in Polyzwitterion/Polyanion Blends

We investigated blends of poly(1-(3-sulfonatopropyl)-2-vinylpyridinium) (P2VPPS) and poly(lithium (trifluoromethane)sulfonimide methacrylate) (poly(MTFSI)Li) at varying molar ratios to gain a mechanistic understanding of ionic conductivity in a miscible polyzwitterion/polyanion system. This dataset contains the raw numerical data corresponding to the figures in the manuscript. The data files include the following information: (1) Experimental Data – includes X-ray and neutron scattering measurements, broadband dielectric spectroscopy (BDS) data, extracted DC conductivity values, differential scanning calorimetry (DSC) and thermogravimetric analysis (TGA) results, extracted glass transition temperatures, etc. (2) CGMD Data – includes molecular dynamics (MD) trajectory files and computed structural correlations. All data files are organized/named according to the figure numbers in the manuscript.

36 MATERIALS SCIENCE

A simulation framework for evaluating electronic order workflows in integrated health records

Electronic health record (EHR) systems are critical to modern healthcare delivery, yet the dynamic workflows that govern electronic order processing remain underexplored. Inefficiencies in these digital pathways can cause delays in care, repetitive workloads, and even patient harm. This study presents a discrete-event simulation framework used to reconstruct and evaluate EHR-based order workflows in a large integrated healthcare system. Using real-world data extracted from the Veterans Health Administration’s Corporate Data Warehouse, the authors mapped order events to standardized state transitions and modeled their progression across different facilities of varying complexity levels. After being calibrated with empirical distributions of transition times and validated against observed time-in-system metrics, the simulation demonstrates close alignment with historical performance. Scenario analyses reveal that resource capacity constraints significantly amplify the impact of electronic order surges, which are reflected in the disproportionate growth in backlogs and processing delays. Adjustments in transition probabilities further increased recirculation and extended workflow paths. Network-based analysis identified Reserved, InProgress, and Completed as structurally critical states that function as hubs within the process network but the transitions in-between also act as major bottlenecks. These results showcased the effectiveness of simulation-based approaches in monitoring EHR order processing performance and evaluating consequences of workflow changes on healthcare network resources planning. The proposed simulation framework provides a scalable data-driven tool to support operational decision-making and improve the efficiency of electronic order management in complex healthcare environments.

Engineering

A Data Processing Pipeline To Extract A Knowledge Graph From Heterogeneous Data For Socio-technical Analysis Of Critical Infrastructure Influence

The code is written in Python and consists of the following pipeline that is implemented in Apache Airflow. This pipeline intends to understand the companies that are directly or indirectly involved with a type of critical infrastructure system at some point in that system's lifecycle. The pipeline takes a configuration file that specifies a list of initial companies to consider, a geographic region of interest, and a set of SEC form types as well as other data sources (e.g. CrunchBase) from which to extract entities and relations. There are four main components to this pipeline as currently implemented: Entity Extraction, Network Construction, Analysis, and Visualization. First, Entity Extraction, is implemented as the `topear-extract_organizations` Apache Airflow workflow. Given an initial query that specifies a geographic region of interest and a time interval, the software will extract CI facilities of interest and organizations that have a direct influence relationship to those facilities (e.g. ownership). During the course of the LDRD, we focused on Electric Vehicle charging stations and this information is available via the Department of Energy (DOE) database on fueling stations maintained by NREL. Within the context of the DOE CESER project, we have focused on Battery Energy Storage Systems (BESS). Second, the Network Extraction component will iteratively construct a social network graph given the set of organizations and people extracted in the previous step. Organizations (and eventually People if desired) are then fed as a query to the `topgear-construct_social_network` Apache Airflow workflow which given a set of initial companies and data sets (e.g. SEC EDGAR form types, OpenCorporates, Crunchbase). This Airflow workflow will iteratively query such data sources to discover relationships with new organizations and people. For example, this module can iteratively query SEC EDGAR for metadata that documents the number of each type of form for the given set of companies and their location. This forms metadata represents a catalog of data sources from SEC EDGAR for the extracted social network knowledge graph. The pipeline then downloads these forms from the website and saves them in a build directory for further processing. These documents are then parsed for entities and relations. Again, we note that in additional to SEC data sources, this step can also pull in information on organizations via API services such as CrunchBase and OpenCorporates or bulk data sources. At the end of this step, the resultant social network, the Critical Infrastructure network, and the edges that encode relationships between organizations and CI facilities, form the Adversarial Socio-Technical Network (ASTN) that informs the analysis. Third, the Analysis component processes these generated ASTN. Previously, that has included the ability to compare prevalence of different vendors for a given infrastructure component type across different regions as well as identify common public and private investors across those vendors. This was demonstrated for EV Charging Stations across several different metropolitan areas within an IEEE PES GridEdge publication. More recently, we have looked at ways to identify infrastructure owners and operators of BESS with the most nameplate capacity across different states as well as other indictors of risk resulting from changes in ownership over time. Finally, the Visualization component consists of an HTML/CSS/JS framework by which users can interact geospatial, operational, and organizational relationships across a given portfolio of Critical Infrastructure facilities. The objective is to provide a library of UI/UX modules that can be repurposed for stakeholder-specific dashboards. All of the modules are related via a common event model that enables UI actions in one view to percolate across the other views.

Weaver, Gabriel [Idaho National Laboratory (INL),

Measurement of the differential cross section for neutral pion production in charged-current muon neutrino interactions on argon with the MicroBooNE detector

We present a measurement of neutral pion production in charged-current interactions using data recorded with the MicroBooNE detector exposed to Fermilab’s booster neutrino beam. The signal comprises one muon, one neutral pion, any number of nucleons, and no charged pions. Studying neutral pion production in the MicroBooNE detector provides an opportunity to better understand neutrino-argon interactions, and is crucial for future accelerator-based neutrino oscillation experiments. Using a dataset corresponding to 6.86 ×10 20 protons on target, we present single-differential cross sections in muon and neutral pion momenta, scattering angles with respect to the beam for the outgoing muon and neutral pion, as well as the opening angle between the muon and neutral pion. Data extracted cross sections are compared to generator predictions. We report good agreement between the data and the models for scattering angles, except for an over-prediction by generators at muon forward angles. Similarly, the agreement between data and the models as a function of momentum is good, except for an underprediction by generators in the medium momentum ranges, 200–400 MeV for muons and 100–200 MeV for pions.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND

Challenges and Vision for Standardization of Biopolymer Data Sets for Machine Learning

Machine learning (ML) is transforming materials research, yet potential for biopolymer discovery remains constrained by fragmented data and nonstandardized reporting. Biopolymers differ significantly from synthetic polymers, requiring specialized approaches to represent their biosynthetic origins, hierarchical structures, and application-specific metrics. In this Perspective, we identify three core challenges limiting biopolymer representation: information encoding, data quality, and data sharing. We describe the most pressing issues and propose commensurate approaches to address each key challenge. Recommendations include the design and adoption of biopolymer-specific fingerprinting and representation frameworks, development of hybrid human-large language model (LLM) data extraction strategies, and expanding Findable, Accessible, Interoperable, Reusable (FAIR)-compliant repositories. We propose a robust foundation to define interoperable, high-quality data sets that capture the full context of biopolymer materials. Standardized metadata, shared ontologies, and community-driven infrastructure would enable scalable, reproducible workflows and accelerate the ML-driven development of biopolymers.

36 MATERIALS SCIENCE

Pore2Chip: All-in-one python tool for soil microstructure analysis and micromodel design

The Pore2Chip Python package is designed to create 2D micromodels using extracted data from 3D X-ray computed tomography (XCT) images. This package helps analyze soil structure and function, allowing for the investigation of hydro-biogeochemical processes that impact mineral extraction and reactivity, oxygen concentrations, and nutrient availability in disturbed or managed soils. Key metrics encompass pore size distributions, pore throat size distributions, and connectivity (pore coordination numbers). The final output is a 2D scalable SVG design representing a core or aggregate. Designs can be fabricated with methods such as laser etching, 3D printing, and photolithography.

lab-on-chip

Deep Design Data Portal (D3P) v0.01

The Deep Design Data Portal (D3P) tool was developed to demonstrate how readily accessible data sources, such as building energy model reports for design and baseline energy performance data for projects, can provide the data required for reporting to an industry initiative (AIA 2030 commitment), as well as more detailed data that makes the industry dataset more valuable to all stakeholders, enabling project level analysis and analysis of BEM industry trends. D3P provides an easier and less time-consuming way for firms to auto-extract data from this data source, compared to the current reporting workflows of the firms. The BEM reports are the first of several data sources that D3P could integrate. D3P also provides the ability for firms to review, compare, and evaluate the performance of their projects to not only their portfolio, but also to the larger anonymized industry dataset created each time a project is added to D3P. The intent of D3P is to become part of a data-sharing ecosystem to assist creating large anonymized industry datasets that are accessible to industry.

Regnier, Cynthia [Lawrence Berkeley National Labor