Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “scientific literature”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Challenges and Advances in Information Extraction from Scientific Literature: a Review

Scientific articles have long been the primary means of disseminating scientific discoveries. Over the centuries, valuable data and potentially groundbreaking insights have been collected and buried deep in the mountain of publications. In materials engineering, such data are spread across technical handbooks specification sheets, journal articles, and laboratory notebooks in myriad formats. Extracting information from papers on a large scale has been a tedious and time-consuming job to which few researchers have wanted to devote their limited time and effort, yet is an activity that is essential for modern data-driven design practices. However, in recent years, significant progress has been made by the computer science community on techniques for automated information extraction from free text. Yet, transformative application of these techniques to scientific literature remains elusive-due not to a lack of interest or effort but to technical and logistical challenges. Using the challenges in the materials science literature as a driving motivation, we review the gaps between state-of-the-art information extraction methods and the practical application of such methods to scientific texts, and offer a comprehensive overview of work that can be undertaken to close these gaps.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Unleashing the Power of Knowledge Extraction from Scientific Literature in Catalysis

Valuable knowledge of catalysis is often hidden in a large amount of scientific literature. There is an urgent need to extract useful knowledge to facilitate scientific discovery. Here this work takes the first step toward the goal in the field of catalysis. Specifically, we construct the first information extraction benchmark data set that covers the field of catalysis and also develop a general extraction framework that can accurately extract catalysis-related entities from scientific literature with 90% extraction accuracy. We further demonstrate the feasibility of leveraging the extracted knowledge to help users better access relevant information in catalysis through an entity-aware search engine and a correlation analysis system.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Advances in scientific literature mining for interpreting materials characterization

Abstract Using synchrotron light sources, such as the National Synchrotron Light Source II at Brookhaven National Laboratory, scientists in fields as diverse as physics, biology, and materials science, identify the atomic structure, chemical composition, or other important properties of varied specimens. x-ray spectroscopy from light sources is particularly valuable for materials research with vast information available about reference spectra in the scientific literature. However, as the technique is applicable to many science domains, searching for information about select x-ray spectroscopy spectra is impeded by the sheer number of publications. Moreover, useful information about the context of an experiment or figures presented in papers can be buried among the details, which takes time to assess. This work presents a scientific literature mining system that supports data acquisition, information extraction, and user interaction for referencing x-ray spectra identification and spectral interpretation. The goal is to provide efficient access to useful spectral data to researchers who may spend only a few days at a synchrotron light source. With this system, users browse a classification tree for papers arranged according to x-ray spectroscopic methods, chemical elements, and x-ray absorption spectroscopy edges. Relevant figures are extracted with sentences from the paper that explain them, known as ‘figure explanatory text.’ Notably, this system focuses on semantic aspects (logical analysis) to find figure explanatory text using deep contextualized word embeddings techniques and contains an interface to obtain labeled data from domain experts that is used to evaluate and improve the model.

Park, Gilchan (ORCID:0000000201536646)↗

A database of thermally activated delayed fluorescent molecules auto-generated from scientific literature with ChemDataExtractor

A database of thermally activated delayed fluorescent (TADF) molecules was automatically generated from the scientific literature. It consists of 25,482 data records with an overall precision of 82%. Among these, 5,349 records have chemical names in the form of SMILES strings which are represented with 91% accuracy; these are grouped in a subsidiary database. Each data record contains one of the following four properties: maximum emission wavelength (λ EM ), photoluminescence quantum yield (PLQY), singlet-triplet energy splitting (ΔE ST ), and delayed lifetime (τ D ). The databases were created through text mining using ChemDataExtractor, a chemistry-aware natural-language-processing toolkit, which has been adapted for TADF research. The text-mined corpus consisted of 2,733 papers from the Royal Society of Chemistry and Elsevier. To the best of our knowledge, these databases are the first databases that have been auto-generated for TADF molecules from existing publications. The databases have been publicly released for experimental and computational applications in the TADF research field.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A thermoelectric materials database auto-generated from the scientific literature using ChemDataExtractor

An auto-generated thermoelectric-materials database is presented, containing 22,805 data records, automatically generated from the scientific literature, spanning 10,641 unique extracted chemical names. Each record contains a chemical entity and one of the seminal thermoelectric properties: thermoelectric figure of merit, ZT; thermal conductivity, κ; Seebeck coefficient, S; electrical conductivity, σ; power factor, PF; each linked to their corresponding recorded temperature, T. The database was auto-generated using the automatic sentence-parsing capabilities of the chemistry-aware, natural language processing toolkit, ChemDataExtractor 2.0, adapted for application in the thermoelectric-materials domain, following a rule-based sentence-simplification step. Data were mined from the text of 60,843 scientific papers that were sourced from three scientific publishers: Elsevier, the Royal Society of Chemistry, and Springer. To the best of our knowledge, this is the first automatically-generated database of thermoelectric materials and their properties from existing literature. The database was evaluated to have a precision of 82.25% and has been made publicly available to facilitate the application of data science in the thermoelectric-materials domain, for analysis, design, and prediction.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Insufficient reporting of x-ray photoelectron spectroscopy instrumental and peak fitting parameters (metadata) in the scientific literature

This study was motivated by earlier observations. It is a systematic examination of the adequacy of reporting of information (metadata) necessary to understand x-ray photoelectron spectroscopy (XPS) data collection and data analysis in the scientific literature. The information for this study was obtained from papers published in three high-quality journals over a six-month period in 2019 and throughout 2021. Each paper was evaluated to determine whether the authors had reported (percentages of the papers properly providing the information are given in parentheses) the spectrometer (66%), fitting software (15%), x-ray source (40%), pass energy (10%), spot size (5%), synthetic peak shapes in fits (10%), backgrounds in fits (10%), whether the XPS data are shown in the main body of the paper or in the supporting information (or both), and whether fitted or unfitted spectra were shown (80% of published spectra are fit). The Shirley background is the most widely used background in XPS peak fitting. The Al Kα source is the most widely used x-ray source for XPS data collection. CASAXPS is the most widely used fitting program for XPS data analysis. Further, there is good agreement between the results gathered during the two years of our survey. There are some hints the situation may be improving. This study also provides a list of the information/parameters that should be reported when XPS is performed.

47 OTHER INSTRUMENTATION↗

A rule-free workflow for the automated generation of databases from scientific literature

Abstract In recent times, transformer networks have achieved state-of-the-art performance in a wide range of natural language processing tasks. Here we present a workflow based on the fine-tuning of BERT models for different downstream tasks, which results in the automated extraction of structured information from unstructured natural language in scientific literature. Contrary to existing methods for the automated extraction of structured compound-property relations from similar sources, our workflow does not rely on the definition of intricate grammar rules. Hence, it can be adapted to a new task without requiring extensive implementation efforts and knowledge. We test our data-extraction workflow by automatically generating a database for Curie temperatures and one for band gaps. These are then compared with manually curated datasets and with those obtained with a state-of-the-art rule-based method. Furthermore, in order to showcase the practical utility of the automatically extracted data in a material-design workflow, we employ them to construct machine-learning models to predict Curie temperatures and band gaps. In general, we find that, although more noisy, automatically extracted datasets can grow fast in volume and that such volume partially compensates for the inaccuracy in downstream tasks.

36 MATERIALS SCIENCE↗

Precursor recommendation for inorganic synthesis by machine learning materials similarity from scientific literature

Synthesis prediction is a key accelerator for the rapid design of advanced materials. However, determining synthesis variables such as the choice of precursor materials is challenging for inorganic materials because the sequence of reactions during heating is not well understood. In this work, we use a knowledge base of 29,900 solid-state synthesis recipes, text-mined from the scientific literature, to automatically learn which precursors to recommend for the synthesis of a novel target material. The data-driven approach learns chemical similarity of materials and refers the synthesis of a new target to precedent synthesis procedures of similar materials, mimicking human synthesis design. When proposing five precursor sets for each of 2654 unseen test target materials, the recommendation strategy achieves a success rate of at least 82%. Our approach captures decades of heuristic synthesis data in a mathematical form, making it accessible for use in recommendation engines and autonomous laboratories.

36 MATERIALS SCIENCE↗

Dataset of solution-based inorganic materials synthesis procedures extracted from the scientific literature

The development of a materials synthesis route is usually based on heuristics and experience. A possible new approach would be to apply data-driven approaches to learn the patterns of synthesis from past experience and use them to predict the syntheses of novel materials. However, this route is impeded by the lack of a large-scale database of synthesis formulations. In this work, we applied advanced machine learning and natural language processing techniques to construct a dataset of 35,675 solution-based synthesis procedures extracted from the scientific literature. Each procedure contains essential synthesis information including the precursors and target materials, their quantities, and the synthesis actions and corresponding attributes. Every procedure is also augmented with the reaction formula. Through this work, we are making freely available the first large dataset of solution-based inorganic materials synthesis procedures.

36 MATERIALS SCIENCE↗

A Database of Stress-Strain Properties Auto-generated from the Scientific Literature using ChemDataExtractor

Abstract There has been an ongoing need for information-rich databases in the mechanical-engineering domain to aid in data-driven materials science. To address the lack of suitable property databases, this study employs the latest version of the chemistry-aware natural-language-processing (NLP) toolkit, ChemDataExtractor, to automatically curate a comprehensive materials database of key stress-strain properties. The database contains information about materials and their cognate properties: ultimate tensile strength, yield strength, fracture strength, Young’s modulus, and ductility values. 720,308 data records were extracted from the scientific literature and organized into machine-readable databases formats. The extracted data have an overall precision, recall and F-score of 82.03%, 92.13% and 86.79%, respectively. The resulting database has been made publicly available, aiming to facilitate data-driven research and accelerate advancements within the mechanical-engineering domain.

Kumar, Pankaj↗

Extracting Material Property Measurements from Scientific Literature with Limited Annotations

Extracting material property data from scientific text is pivotal for advancing data-driven research in chemistry and materials science; however, the extensive annotation effort required to produce training data for named entity recognition (NER) models for this task often makes it a barrier to extracting specialized data sets. Here, in this work, we present a comparative study of the conventional, supervised NER methodology to alternative few-shot learning architectures and large language model (LLM)-based approaches that mitigate the need to label large training data sets. We find that the best-performing LLM (GPT-4o) not only excels in directly extracting relevant material properties based on limited examples but also enhances supervised learning through data augmentation. We supplement our findings with error and data quality assessments to provide a nuanced understanding of factors that impact property measurement extraction.

36 MATERIALS SCIENCE↗

PhysBERT: A text embedding model for physics scientific literature

The specialized language and complex concepts in physics pose significant challenges for information extraction through Natural Language Processing (NLP). Central to effective NLP applications is the text embedding model, which converts text into dense vector representations for efficient information retrieval and semantic analysis. In this work, we introduce PhysBERT, the first physics-specific text embedding model. Pre-trained on a curated corpus of 1.2 × 106 arXiv physics papers and fine-tuned with supervised data, PhysBERT outperforms leading general-purpose models on physics-specific tasks, including the effectiveness in fine-tuning for specific physics subdomains.

Hellert, Thorsten (ORCID:0000000227970926)↗

FASAC Technical Assessment Report: Soviet Space Science Research

This report is the work of a panel of eight US scientists who surveyed and assessed Soviet research in the spare sciences. All of the panelists were very familiar with Soviet research through their knowledge of the published scientific literature and personal contacts with Soviet and other foreign colleagues. In addition, all of the panelists reviewed considerable additional open literature--scientific, and popular, including news releases. The specific disciplines of Soviet space science research examined in detail for the report were: solar-terrestrial research, lunar and planetary research, space astronomy and astrophysics, and, life sciences. The Soviet Union has in the past carried out an ambitious program in lunar exploration and, more recently, in studies of the inner planets, Mars and especially Venus. The Soviets have provided scientific data about the latter planet which has been crucial for studies of the planet's evolution. Future programs envision an encounter with Halley's Comet, in March 1986, and missions to Mars and asteroids. The Soviet programs in the life sciences and solar-terrestrial research have been long-lasting and systematically pursued. Much of the ground-based and space-based research in these two disciplines appears to be motivated by the requirement to establish long-term human habitation in near-Earth space. The Soviet contributions to new discoveries and understanding in observational space astronomy and astrophysics have been few. This is in significant contrast to the very excellent theoretical work contributed by Soviet scientists in this discipline.

Lanzerotti, L. J.↗

Cognitive analysis of metabolomics data for systems biology

Cognitive computing is revolutionizing the way big data are processed and integrated, with artificial intelligence (AI) natural language processing (NLP) platforms helping researchers to efficiently search and digest the vast scientific literature. Most available platforms have been developed for biomedical researchers, but new NLP tools are emerging for biologists in other fields and an important example is metabolomics. NLP provides literature-based contextualization of metabolic features that decreases the time and expert-level subject knowledge required during the prioritization, identification and interpretation steps in the metabolomics data analysis pipeline. Here, we describe and demonstrate four workflows that combine metabolomics data with NLP-based literature searches of scientific databases to aid in the analysis of metabolomics data and their biological interpretation. Additionally, the four procedures can be used in isolation or consecutively, depending on the research questions. The first, used for initial metabolite annotation and prioritization, creates a list of metabolites that would be interesting for follow-up. The second workflow finds literature evidence of the activity of metabolites and metabolic pathways in governing the biological condition on a systems biology level. The third is used to identify candidate biomarkers, and the fourth looks for metabolic conditions or drug-repurposing targets that the two diseases have in common. The protocol can take 1–4 h or more to complete, depending on the processing time of the various software used.

59 BASIC BIOLOGICAL SCIENCES↗

Development of techniques for the in situ observation of OH and HO2 for studies of the impact of high-altitude supersonic aircraft on the stratosphere

This three-year project supported the construction, calibration, and deployment of a new instrument to measure the OH and HO2 radicals on the NASA ER-2 aircraft. The instrument has met and exceeded all of its design goals. The instrumentation represents a true quantum leap in performance over that achieved in previous HO(x) instruments built in our group. Sensitivity for OH was enhanced by over two orders of magnitude as the weight fell from approximately 1500 to less than 200 Kg. Reliability has been very high: HO(x) data are available for all flights during the first operational mission, the Stratospheric Photochemistry, Aerosols, and Dynamics Expedition (SPADE). The results of that experiment have been reported in the scientific literature and at conferences. Additionally, measurements of H2O and O3 were made and have been reported in the scientific literature. The measurements demonstrated the important role that OH and HO2 play in determining the concentration of ozone in the lower stratosphere. During the SPADE, campaign the measurements demonstrated that the catalytic removal is dominated by processes involving the odd-hydrogen and halogen radicals-and extremely important constraint for photochemical models that are being used to assess the potential deleterious effects of super-sonic aircraft effluent on the burden of stratospheric ozone.

Anderson, James G.↗

Development of techniques for the In Situ observation of OH and HO2 for studies of the impact of high-altitude supersonic aircraft on the stratosphere

This three-year project supported the construction, calibration, and deployment of a new instrument to measure the OH and HO2 radicals on the NASA Er-2 aircraft. The instrument has met and exceeded all of its design goals. The instrumentation represents a true quantum leap in performance over that achieved in previous HO(x) instruments built in our group. Sensitivity of OH was enhanced by over two orders of magnitude as the weight fell from approximately 1500 to less than 200 Kg. Reliability has been very high: HO(x) data are available for all flights during the first operational mission, the Stratospheric Photochemistry, aerosols, and Dynamics Expedition (SPADE). The results of that experiment have been reported in the scientific literature and at conferences. Additionally, measurements of H2O and O3 were made and have been reported in the scientific literature. The measurements demonstrate the important role that OH and HO2 play in determining the concentration of ozone in the lower stratosphere. During the SPADE campaign, the measurements demonstrate that the catalytic removal is dominated by processes involving the odd-hydrogen and halogen radicals-an extremely important constraint for photochemical models that are being used to assess the potential deleterious effects of super-sonic aircraft effluent on the burden of stratospheric ozone.

Anderson, James G.↗