Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Scientific literature”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

A Database of Stress-Strain Properties Auto-generated from the Scientific Literature using ChemDataExtractor

Abstract There has been an ongoing need for information-rich databases in the mechanical-engineering domain to aid in data-driven materials science. To address the lack of suitable property databases, this study employs the latest version of the chemistry-aware natural-language-processing (NLP) toolkit, ChemDataExtractor, to automatically curate a comprehensive materials database of key stress-strain properties. The database contains information about materials and their cognate properties: ultimate tensile strength, yield strength, fracture strength, Young’s modulus, and ductility values. 720,308 data records were extracted from the scientific literature and organized into machine-readable databases formats. The extracted data have an overall precision, recall and F-score of 82.03%, 92.13% and 86.79%, respectively. The resulting database has been made publicly available, aiming to facilitate data-driven research and accelerate advancements within the mechanical-engineering domain.

Kumar, Pankaj

Extracting Material Property Measurements from Scientific Literature with Limited Annotations

Extracting material property data from scientific text is pivotal for advancing data-driven research in chemistry and materials science; however, the extensive annotation effort required to produce training data for named entity recognition (NER) models for this task often makes it a barrier to extracting specialized data sets. Here, in this work, we present a comparative study of the conventional, supervised NER methodology to alternative few-shot learning architectures and large language model (LLM)-based approaches that mitigate the need to label large training data sets. We find that the best-performing LLM (GPT-4o) not only excels in directly extracting relevant material properties based on limited examples but also enhances supervised learning through data augmentation. We supplement our findings with error and data quality assessments to provide a nuanced understanding of factors that impact property measurement extraction.

36 MATERIALS SCIENCE

PhysBERT: A text embedding model for physics scientific literature

The specialized language and complex concepts in physics pose significant challenges for information extraction through Natural Language Processing (NLP). Central to effective NLP applications is the text embedding model, which converts text into dense vector representations for efficient information retrieval and semantic analysis. In this work, we introduce PhysBERT, the first physics-specific text embedding model. Pre-trained on a curated corpus of 1.2 × 106 arXiv physics papers and fine-tuned with supervised data, PhysBERT outperforms leading general-purpose models on physics-specific tasks, including the effectiveness in fine-tuning for specific physics subdomains.

Hellert, Thorsten (ORCID:0000000227970926)

Openpronghorn

OpenPronghorn is a simulation tool specifically tailored for modeling thermal-hydraulic phenomena in advanced nuclear reactors. It is built on the Multiphysics Object-Oriented Simulation Environment (MOOSE), an open-source platform that facilitates the development of high-performance scientific computing applications. OpenPronghorn solves the Navier-Stokes equations, which describe the conservation of mass, momentum, and energy in fluid flows, using the finite volume numerical method. The code supports a wide range of fluid flow conditions that are applicable to nuclear reactors, including incompressible and weakly compressible flows, as well as single-phase and multiphase flows. It is capable of modeling diverse flow regimes, including laminar and turbulent flows, using various turbulence models such as the standard k-epsilon models, the v2f model, and the mixing length model. For multiphase flows, OpenPronghorn employs a mixture a Eulerian modeling approach with mixture, drift-flux, and full Eulerian models, and includes open-sourced interfacial transfer correlations for drag, exchange, and heat transfer coming from the scientific literature. OpenPronghorn's modular design allows it to handle multiscale simulations, ranging from detailed Reynolds-Averaged Navier Stokes (RANS) simulations to coarse-mesh and lumped parameter models. This flexibility enables users to perform high-fidelity simulations of specific reactor components as well as system-level analyses of entire reactor circuits. The code can be coupled with other MOOSE-based tools using the MultiApp system, allowing for the transfer of coupling quantities such as mass flow rates, heat fluxes, and boundary conditions between different simulation scales. One of the main features of OpenPronghorn is the it includes built-in validation cases from the open-source scientific literature and supports the implementation of user-defined models and correlations through MOOSE's FunctorMaterial system. OpenPronghorn is designed to be computationally efficient, leveraging the SIMPLE projection method for large-scale problems, and can be run on high-performance computing systems to handle the extensive computational demands of detailed reactor simulations. Overall, OpenPronghorn is a versatile and robust tool that provides critical insights into the thermal-hydraulic behavior of advanced nuclear reactors, supporting the design, safety, and optimization of next-generation nuclear energy systems.

Retamales, Mauricio Eduardo Tano [Idaho National L

Trends and 2025 Insights on the Rise of Electric Vehicles in the USA

Plug-in electric vehicles (EVs) are reshaping the transportation energy landscape, providing a practical alternative to petroleum fuels for a growing number of applications. EV sales grew 55x in the past decade (2014-2024) and 6x since 2020, driven by technological progress enabled by policies to reduce transportation emissions as well as industrial plans motivated by strategic value of EVs for global competitiveness, jobs and geopolitics. In 2024, 22% of passenger cars sold globally were EVs and opportunities for EVs beyond on-road applications are growing, including solutions to electrify off-road vehicles, maritime and aviation. This Review updates and expands our 2020 assessment of the scientific literature and describes the current status and future projections of EV markets, charging infrastructures, vehicle-grid integration and supply chains in the USA. EV is the lowest-emission motorized on-road transportation option, with life-cycle emissions decreasing as electricity emissions continue to decrease. Charging infrastructure grew in line with EV adoption but providing ubiquitous reliable and convenient charging remains a challenge. EVs are reducing electricity costs in several US markets and coordinated EV charging can improve grid resilience and reduce electricity costs for all consumers. The current trajectory of technology improvement and industrial investments points to continued acceleration of EVs.

33 ADVANCED PROPULSION SYSTEMS

Expert evaluation of LLM world models: A high-T c superconductivity case study

Large Language Models (LLMs) show great promise as a powerful tool for scientific literature exploration. However, their effectiveness in providing scientifically accurate and comprehensive answers to complex questions within specialized domains remains an active area of research. Using the field of high-temperature cuprates as an exemplar, we evaluate the ability of LLM systems to understand the literature at the level of an expert. We construct an expert-curated database of 1,726 scientific papers that covers the history of the field, and a set of 67 expert-formulated questions that probe deep understanding of the literature. We then evaluate six different LLM-based systems for answering these questions, including both commercially available closed models and a custom retrieval-augmented generation (RAG) system capable of retrieving images alongside text. Experts then evaluate the answers of these systems against a rubric that assesses balanced perspectives, factual comprehensiveness, succinctness, and evidentiary support. Among the six systems, two using RAG on curated literature outperformed existing closed models across key metrics, particularly in providing comprehensive and well-supported answers. We discuss promising aspects of LLM performances as well as critical short-comings of all the models. The set of expert-formulated questions and the rubric will be valuable for assessing expert level performance of LLM based reasoning systems.

36 MATERIALS SCIENCE

MaTableGPT: GPT‐Based Table Data Extractor from Materials Science Literature

Abstract Efficiently extracting data from tables in the scientific literature is pivotal for building large‐scale databases. However, the tables reported in materials science papers exist in highly diverse forms; thus, rule‐based extractions are an ineffective approach. To overcome this challenge, the study presents MaTableGPT, which is a GPT‐based table data extractor from the materials science literature. MaTableGPT features key strategies of table data representation and table splitting for better GPT comprehension and filtering hallucinated information through follow‐up questions. When applied to a vast volume of water splitting catalysis literature, MaTableGPT achieves an extraction accuracy (total F1 score) of up to 96.8%. Through comprehensive evaluations of the GPT usage cost, labeling cost, and extraction accuracy for the learning methods of zero‐shot, few‐shot, and fine‐tuning, the study presents a Pareto‐front mapping where the few‐shot learning method is found to be the most balanced solution owing to both its high extraction accuracy (total F1 score >95%) and low cost (GPT usage cost of 5.97 US dollars and labeling cost of 10 I/O paired examples). The statistical analyses conducted on the database generated by MaTableGPT revealed valuable insights into the distribution of the overpotential and elemental utilization across the reported catalysts in the water splitting literature.

Yi, Gyeong Hoon [Computational Science Research Ce

poppler-science

The “Poppler-science” software is a fork of the existing open-source Poppler project (https://poppler.freedesktop.org/) for converting PDF files to text. Modifications to the Poppler source code include (a) per-glyph optical character recognition (for correcting the non-standard font glyph remapping that is common in the scientific literature), (b) inference of text markup for commonly used scientific formatting (like superscripts and subscripts), (c) table and figure recognition, and (d) improved ordering of text output for complex scientific manuscripts (e.g., multi-column text, figure and table captions, etc.).

Gans, Jason [Los Alamos National Laboratory]

WHOLESCALE - Water & Hole Observations Leverage Effective Stress Calculations And Lessen Expenses (Final Technical Report 2020 - 2024)

The WHOLESCALE acronym stands for Water & Hole Observations Leverage Effective Stress Calculations and Lessen Expenses. The goal of the WHOLESCALE project is to simulate the spatial distribution and temporal evolution of stress in the geothermal system at San Emidio in Nevada, United States. To reach this goal, the WHOLESCALE team has developed a methodology to incorporate and interpret data from four methods of measurement into a multi-physics model that couples thermal, hydrological, and mechanical (T H-M) processes. The WHOLESCALE team has applied this methodology at the San Emidio geothermal field, located ~100 km north of Reno, Nevada in the northwestern Basin and Range province. The WHOLESCALE team includes 30 individuals working at two universities, two national laboratories, and one industry partner. Two master-degree students and five post-doctoral researchers have gained professional experience and earned partial financial support via the WHOLESCALE project. The WHOLESCALE team has taken advantage of the perturbations created by changes in pumping operations during planned shutdowns in 2016, 2021, and 2022 to infer temporal changes in the state of stress in the geothermal system at San Emidio, Nevada, U.S. The WHOLESCALE results support the working hypothesis that increasing pore-fluid pressure reduces the effective normal stress acting across fault zones. During normal operations, pumping in deep production wells decreases fluid pressures and thus increases the effective normal stresses on faults, reducing microseismicity. During planned shutdowns, the cessation of production increases pore-fluid pressure and reduces effective normal stress. The WHOLESCALE products generated during the 4-year period between 2020 and 2024 include: three articles published in the open-access, peer-reviewed scientific literature, two master’s theses, 20 presentations or papers at scientific conferences, and 17 data sets available on public repositories. The WHOLESCALE project has been completed in two phases that included three performance periods separated by two Go/No-go Stage Gate Reviews. Tasks were classified by data type (i.e., Geologic Structure, Borehole, Geodesy, Hydrology, Seismology, and Modeling). The first phase of the project started July 31, 2020 and included ongoing project coordination (Task 1), a project kickoff (Task 2), analysis of existing data (Task 3), development of the initial stress model & deployment design (Task 4), and Go/No-go Decision Point #1 (Task 5). Phase II began with implementing the 2022 deployment (Task 6), followed by Go/No-go Decision Point #2 (Task 7) The remainder of Phase II consisted of analyzing data collected during deployment (Task 8), calibration of the stress model on all observations (Task 9), and the Final Review (August 23, 2024) & Reporting (Task 10).

15 GEOTHERMAL ENERGY

Acquisition of absorption and fluorescence spectral data using chatbots

Spectra – the lifeblood of photochemistry – have been very difficult to find in the literature. Chatbots, remarkably, may enable their more efficient acquisition and prove to be generally powerful tools for searching the scientific literature.

Taniguchi, Masahiko [Department of Chemistry, Nort

Literature Review of Selected Publications Relevant to Carbon Dioxide Pipeline and Storage Systems

The annotated bibliographies provided in this document focus on identifying and reviewing environmental impact statements (EIS), environmental assessments (EA), and scientific literature relevant to aspects of CO 2 pipeline and storage construction and operation. The purpose of these summaries is to aid stakeholders responsible for the preparation of documents compliant with the National Environmental Policy Act (NEPA) requirements to have access to a quick and comprehensive guide of the literature that can inform their activities. They aim to inform the development of EISs and EAs by identifying potential environmental concerns and mitigation strategies. The annotated bibliographies captured in this document cover the general environmental assessment and impact topics relevant to CO 2 pipeline and storage construction and operation activities, with the subject of waste creation and handling being specifically pulled out into its own section. The reasoning behind a dedicated section for waste is that many EAs and EISs focus on the description of waste and its impacts, be it caused in routine operation or as a result of an accident.

54 ENVIRONMENTAL SCIENCES

AI-Powered Knowledge Graphs for Neuromorphic and Energy-Efficient Computing

The surge in scientific literature obscures breakthroughs and hinders the discovery of new research paths. We propose an artificial intelligence (AI) powered framework using large language models (LLMs) and knowledge graphs (KGs) to automate parts of scientific discovery, focusing on energy-efficient AI circuits. Our hybrid approach combines LLMs, structured data, and ontology-based reasoning to construct a comprehensive knowledge graph that integrates insights across computational neuroscience, spiking neuron models, learning rules, architectural motifs, and neuromorphic device technologies. This multi-domain representation enables the generation of hypotheses that connect biological function with implementable, energy-efficient hardware architectures. Using KG embeddings and graph neural networks, the framework generates hypotheses for novel circuits, validates them through optimization on exascale HPC systems, and with tools like SuperNeuro and Fugu, the most promising designs will be prototyped in hardware. This open-source system aims to accelerate discoveries and bridging neuroscience with hardware innovation, drive collaboration, and unlock new opportunities in low-power AI computing.

Gautam, Ashish [ORNL]

Gender in Mineral Names

Minerals are the fundamental constituents of Earth, and mineral names appear in scientific literature for disciplines including geology, chemistry, materials science, biology, and medicine, among others. Choosing a name is the full responsibility of the authors of new mineral proposals submitted to the International Mineralogical Association (IMA). Scientific nomenclature and its traditions have evolved over time and, consequently, mineral names track changes in the landscape of mineralogy with respect to language, technology, and culture. To evaluate these changes, the namesake information for all 5896 minerals approved by the IMA or ‘grandfathered’ into use as of December 2022 was recorded and categorized within a workable database. The compiled information yields diverse insights into the intersection of science and culture and could also be used to project future trends. In this study, we used the name database to investigate gender diversity among mineral eponyms. More than half (c. 54%) of all mineral species are named after people, the identities of whom are largely a reflection of the people that have historically been involved, in one way or another, in the geosciences and in the mining industry. Of the 2738 people with minerals named for them, approximately 6.1% are (interpreted to be) women. Nearly all minerals named for women were named during the last sixty years, although the rate of growth in the year-on-year percentage of women among new mineral namesakes has slowed since about 1985. If current and historical trends hold, our model predicts that women will not comprise more than about 10.35% of newly established mineral namesakes in future years. The representation of women among mineral namesakes also differs starkly among countries. For example, Russians comprise 43.11% of women with minerals named for them, but account for only 15.12% of all eponyms. However, there are additional disparities beyond the proportions of namesakes. For scientists who were alive when a mineral was named for them, women were an average of 3.74 years older than men when evaluated over the same timespan (1954–2022). These results demonstrate that gender-based disparities are imprinted into current mineral nomenclature and indicate that gender parity among new mineral namesakes is impossible without unprecedented changes in the upstream demographics that are most likely to affect naming trends.

58 GEOSCIENCES

32 examples of LLM applications in materials science and chemistry: towards automation, assistants, agents, and accelerated scientific discovery

Abstract Large language models (LLMs) are reshaping many aspects of materials science and chemistry research, enabling advances in molecular property prediction, materials design, scientific automation, knowledge extraction, and more. Recent developments demonstrate that the latest class of models are able to integrate structured and unstructured data, assist in hypothesis generation, and streamline research workflows. To explore the frontier of LLM capabilities across the research lifecycle, we review applications of LLMs through 32 total projects developed during the second annual LLM hackathon for applications in materials science and chemistry, a global hybrid event. These projects spanned seven key research areas: (1) molecular and material property prediction, (2) molecular and material design, (3) automation and novel interfaces, (4) scientific communication and education, (5) research data management and automation, (6) hypothesis generation and evaluation, and (7) knowledge extraction and reasoning from the scientific literature. Collectively, these applications illustrate how LLMs serve as versatile predictive models, platforms for rapid prototyping of domain-specific tools, and much more. In particular, improvements in both open source and proprietary LLM performance through the addition of reasoning, additional training data, and new techniques have expanded effectiveness, particularly in low-data environments and interdisciplinary research. As LLMs continue to improve, their integration into scientific workflows presents both new opportunities and new challenges, requiring ongoing exploration, continued refinement, and further research to address reliability, interpretability, and reproducibility.

Computer Science

A Review of Abrupt Permafrost Thaw: Definitions, Usage, and a Proposed Conceptual Framework

Purpose of ReviewWe review how ‘abrupt thaw’ has been used in published studies, compare these definitions to abrupt processes in other Earth science disciplines, and provide a definitive framework for how abrupt thaw should be used in the context of permafrost science.Recent FindingsWe address several aspects of permafrost systems necessary for abrupt thaw to occur and propose a framework for classifying permafrost processes as abrupt thaw in the future. Based on a literature review and our collective expertise, we propose that abrupt thaw refers to thaw processes that lead to a substantial persistent environmental change within a few decades. Abrupt thaw typically occurs in ice-rich permafrost but may be initiated in ice-poor permafrost by external factors such as hydrologic change (i.e., increased streamflow, soil moisture fluctuations, altered groundwater recharge) or wildfire.SummaryPermafrost thaw alters greenhouse gas emissions, soil and vegetation properties, and hydrologic flow, threatening infrastructure and the cultures and livelihoods of northern communities. The term ‘abrupt thaw’ has emerged in scientific discourse over the past two decades to differentiate processes that rapidly impact large depths of permafrost, such as thermokarst, from more gradual, top-down thaw processes that impact centimeters of near-surface permafrost over years to decades. However, there has been no formal definition for abrupt thaw and its use in the scientific literature has varied considerably. Our standardized definition of abrupt thaw offers a path forward to better understand drivers and patterns of abrupt thaw and its consequences for global greenhouse gas budgets, impacts to infrastructure and land-use, and Arctic policy- and decision-making.

Webb, Hailey

What Is a Polyolefin? A Critical Overview of Ethylene Copolymers Used as Solar Photovoltaic Module Encapsulants

In recent years, photovoltaic (PV) encapsulant films marketed as polyolefins (POs), more specifically as PO elastomers (POEs) and thermoplastic POs (TPOs), have gained significant market share and are projected to become the dominant encapsulation films by 2030. Relative to other industries, there are significant misconceptions about the term PO in the PV industry. Both in the scientific literature as well as in sales and advertising, the terms PO, POE, and TPO are often misused to describe the same type of material with comparable properties, while in reality these may each consist of separate material classes. This paper provides a comprehensive literature and market review, to showcase a broad range of PO and other ethylene copolymer encapsulants from recent studies, and discusses the materials' properties to clarify what constitutes a “polyolefin.” In addition, to promote a clearer comparison of encapsulant properties, we propose a two‐dimensional taxonomy to categorize polymers used in module manufacturing, including POs. In terms of improving the reliability of solar PV modules, PO‐based encapsulants have several advantages (including lower water uptake and ion diffusion), but might come with disadvantages too, such as a more complex processing and a higher sensitivity to the storage conditions and shelf life. All this might prospectively impact adhesion properties of the encapsulant to other materials' interfaces (glass, cells etc.) and end‐product quality. Because the track record of field‐deployed PV modules containing PO encapsulants is also limited, we hope to contribute to better material understanding and precision in communication in PV to secure quality.

14 SOLAR ENERGY

Text Mining for Process–Structure–Properties Relationships in Metals

With the advent of large language models (LLMs), the vast unstructured text within millions of academic papers is increasingly accessible for materials discovery—although significant challenges remain. While LLMs offer promising few- and zero-shot learning capabilities, particularly valuable in the materials domain where expert annotations are scarce, general-purpose LLMs often fail to address key materials-specific queries without further adaptation. To bridge this gap, fine-tuning LLMs on human-labeled data is essential for effective structured knowledge extraction (Liu in The Importance of Human-Labeled Data in the Era of LLMs, 2023). Here, in this study, we introduce a novel annotation schema designed to extract generic process–structure–properties relationships from scientific literature. We demonstrate the utility of this approach using a dataset of 128 abstracts, with annotations drawn from two distinct domains: high-temperature materials (Domain I) and uncertainty quantification in simulating materials microstructure (Domain II). Initially, we developed a conditional random field (CRF) model based on MatBERT—a domain-specific BERT variant—and evaluated its performance on Domain I. Subsequently, we compared this model with a fine-tuned LLM (GPT-4o from OpenAI) under identical conditions. Our results indicate that fine-tuning LLMs can significantly improve entity extraction performance over the BERT-CRF baseline on Domain I. However, when additional examples from Domain II were incorporated, the performance of the BERT-CRF model became comparable to that of the GPT-4o model. These findings underscore the potential of our schema for structured knowledge extraction and highlight the complementary strengths of both modeling approaches.

Materials science

Large language model-driven database for thermoelectric materials

Thermoelectric materials have the ability to convert waste heat into electricity, offering a valuable solution for energy harvesting. However, their widespread use is hindered by low conversion efficiency, the reliance on expensive rare earth elements, and the environmental and regulatory concerns associated with lead-based materials. A fast and cost-effective way to identify highly efficient thermoelectric materials is through data-driven methods. These approaches rely on robust and comprehensive datasets to train models. Although there are several databases on thermoelectric materials, there is still a need to collect and integrate experimental data from peer-reviewed research articles to capture diverse compositions and properties of materials. Here, in this work, we developed a comprehensive database of 7,123 thermoelectric compounds, containing key information such as chemical composition, structural detail, seebeck coefficient, electrical and thermal conductivity, power factor, and figure of merit (ZT). We used the GPTArticleExtractor workflow, powered by large language models (LLM), to extract and curate data automatically from the scientific literature published in Elsevier journals. This process enabled the creation of a structured database that addresses the challenges of manual data collection. The open access database could stimulate data-driven research and advance thermoelectric material analysis and discovery.

Database