Engineering Papers⌕ Search

Engineering topics

Cole, Jacqueline M.

Publications and source records attributed to Cole, Jacqueline M..

How Beneficial Is Pretraining on a Narrow Domain-Specific Corpus for Information Extraction about Photocatalytic Water Splitting?

Language models trained on domain-specific corpora have been employed to increase the performance in specialized tasks. However, little previous work has been reported on how specific a “domain-specific” corpus should be. Here, we test a number of language models trained on varyingly specific corpora by employing them in the task of extracting information from photocatalytic water splitting. We find that more specific corpora can benefit performance on downstream tasks. Furthermore, PhotocatalysisBERT, a pretrained model from scratch on scientific papers on photocatalytic water splitting, demonstrates improved performance over previous work in associating the correct photocatalyst with the correct photocatalytic activity during information extraction, achieving a precision of 60.8(+11.5)% and a recall of 37.2(+4.5)%.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A database of thermally activated delayed fluorescent molecules auto-generated from scientific literature with ChemDataExtractor

A database of thermally activated delayed fluorescent (TADF) molecules was automatically generated from the scientific literature. It consists of 25,482 data records with an overall precision of 82%. Among these, 5,349 records have chemical names in the form of SMILES strings which are represented with 91% accuracy; these are grouped in a subsidiary database. Each data record contains one of the following four properties: maximum emission wavelength (λ EM ), photoluminescence quantum yield (PLQY), singlet-triplet energy splitting (ΔE ST ), and delayed lifetime (τ D ). The databases were created through text mining using ChemDataExtractor, a chemistry-aware natural-language-processing toolkit, which has been adapted for TADF research. The text-mined corpus consisted of 2,733 papers from the Royal Society of Chemistry and Elsevier. To the best of our knowledge, these databases are the first databases that have been auto-generated for TADF molecules from existing publications. The databases have been publicly released for experimental and computational applications in the TADF research field.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Automated Construction of a Photocatalysis Dataset for Water-Splitting Applications

We present an automatically generated dataset of 15,755 records that were extracted from 47,357 papers. These records contain water-splitting activity in the presence of certain photocatalysts, along with additional information about the chemical reaction conditions under which this activity was recorded. These conditions include any co-catalysts and additives that were present during water splitting, the length of time for which the photocatalytic experiment was conducted, and the type of light source used, including its wavelength. Despite the text extraction of such a wide range of chemical reaction attributes, the dataset afforded good precision (71.2%) and recall (36.3%). These figures-of-merit were calculated based on a random sample of open-access papers from the corpus. Mining such a complex set of attributes required the development of novel techniques in knowledge extraction and interdependency resolution, leveraging inter- and intra-sentence relations, which are also described in this paper. We present a new version (version 2.2) of the chemistry-aware text-mining toolkit ChemDataExtractor, in which these new techniques are included.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

In‐Silico Device Performance Prediction of Cosensitizer Dye Pairs for Dye‐Sensitized Solar Cells

Abstract Endeavors in the field of dye‐sensitized solar cells (DSCs) have shown great promise when adopting a data‐driven approach to materials discovery, such as successful molecular‐scale predictions of light‐harvesting chromophores. However, predictions of DSC dyes would become much more sophisticated if a molecular‐to‐macroscopic DSC device prediction methodology existed. Thereby, a fully computational pipeline is presented that predicts device‐performance parameters of DSCs which contain varying dye combinations. Optimal pairing of complementary dyes is identified via a data‐driven workflow that affords cosensitized DSCs with maximum power‐conversion efficiencies. Six high‐performing DSC dyes are paired with partner dyes that are screened from a database of 8488 compounds using sequential heuristic filters. Existing models that predict short‐circuit‐current density ( J SC ) and open‐circuit voltage ( V OC ) parameters are adapted to predict singly sensitized and cosensitized DSC performance. The predictions for J sc values of singly sensitized devices match experimental literature values with comparable accuracy to more computationally costly methods. Five out of six dye pairings are predicted to have greater J SC values when cosensitized compared to their corresponding singly sensitized devices, including two pairs that show strong J sc boosts of +13% and +12% when cosensitized. Thus, the prospect of an entirely in‐silico prediction pipeline for DSC performance that can be used to realize the fully automated design of optimized cosensitized DSCs is demonstrated.

14 SOLAR ENERGY↗

A thermoelectric materials database auto-generated from the scientific literature using ChemDataExtractor

An auto-generated thermoelectric-materials database is presented, containing 22,805 data records, automatically generated from the scientific literature, spanning 10,641 unique extracted chemical names. Each record contains a chemical entity and one of the seminal thermoelectric properties: thermoelectric figure of merit, ZT; thermal conductivity, κ; Seebeck coefficient, S; electrical conductivity, σ; power factor, PF; each linked to their corresponding recorded temperature, T. The database was auto-generated using the automatic sentence-parsing capabilities of the chemistry-aware, natural language processing toolkit, ChemDataExtractor 2.0, adapted for application in the thermoelectric-materials domain, following a rule-based sentence-simplification step. Data were mined from the text of 60,843 scientific papers that were sourced from three scientific publishers: Elsevier, the Royal Society of Chemistry, and Springer. To the best of our knowledge, this is the first automatically-generated database of thermoelectric materials and their properties from existing literature. The database was evaluated to have a precision of 82.25% and has been made publicly available to facilitate the application of data science in the thermoelectric-materials domain, for analysis, design, and prediction.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Perovskite- and Dye-Sensitized Solar-Cell Device Databases Auto-generated Using ChemDataExtractor

The number of scientific publications reporting cutting-edge third-generation photovoltaic devices is increasing rapidly, owing to the pressing need to develop renewable-energy technologies that address the climate-change crisis. Consequently, the field could benefit from a central repository where photovoltaic-performance metrics, such as the power-conversion efficiency (η) are recorded. We present two automatically generated databases that contain photovoltaic properties and device material data for dye-sensitized solar cells (DSCs) and perovskite solar cells (PSCs), totalling 660,881 data entries representing 57,678 photovoltaic devices. The databases were generated by applying the text-mining toolkit ChemDataExtractor on a corpus of 25,720 articles. A multi-faceted evaluation, incorporating manual and automatic methods, was applied to ensure that the data contained therein were of the highest quality, with precision metrics ranging from 73.1% to 95.8%. The DSC database contains 475,045 entries representing 41,680 devices, and the PSC database contains 185,836 entries representing 15,818 devices. The databases are available in MongoDB and JSON formats, which can be queried in Python, R, Java and MATLAB for data-driven photovoltaic materials discovery.

14 SOLAR ENERGY↗

A database of refractive indices and dielectric constants auto-generated using ChemDataExtractor

The ability to auto-generate databases of optical properties holds great potential for advancing optical research, especially with regards to the data-driven discovery of optical materials. An optical property database of refractive indices and dielectric constants is presented, which comprises a total of 49,076 refractive index and 60,804 dielectric constant data records on 11,054 unique chemicals. The database was auto-generated using the state-of-the-art natural language processing software, ChemDataExtractor, using a corpus of 388,461 scientific papers. The data repository offers a representative overview of the information on linear optical properties that resides in scientific papers from the past 30 years. Public availability of these data will enable a quick search for the optical property of certain materials. The large size of this repository will accelerate data-driven research on the design and prediction of optical materials and their properties. To the best of our knowledge, this is the first auto-generated database of optical properties from a large number of scientific papers. We provide a web interface to aid the use of our database.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

PDFDataExtractor: A Tool for Reading Scientific Text and Interpreting Metadata from the Typeset Literature in the Portable Document Format

The layout of portable document format (PDF) files is constant to any screen, and the metadata therein are latent, compared to mark-up languages such as HTML and XML. No semantic tags are usually provided, and a PDF file is not designed to be edited or its data interpreted by software. However, data held in PDF files need to be extracted in order to comply with opensource data requirements that are now government-regulated. In the chemical domain, related chemical and property data also need to be found, and their correlations need to be exploited to enable data science in areas such as data-driven materials discovery. Such relationships may be realized using text-mining software such as the “chemistry-aware” natural-language-processing tool, ChemDataExtractor; however, this tool has limited data-extraction capabilities from PDF files. This study presents the PDFDataExtractor tool, which can act as a plug-in to ChemDataExtractor. It outperforms other PDF-extraction tools for the chemical literature by coupling its functionalities to the chemical-named entityrecognition capabilities of ChemDataExtractor. The intrinsic PDF-reading abilities of ChemDataExtractor are much improved. The system features a template-based architecture. This enables semantic information to be extracted from the PDF files of scientific articles in order to reconstruct the logical structure of articles. While other existing PDF-extracting tools focus on quantity mining, this template-based system is more focused on quality mining on different layouts. PDFDataExtractor outputs information in JSON and plain text, including the metadata of a PDF file, such as paper title, authors, affiliation, email, abstract, keywords, journal, year, document object identifier (DOI), reference, and issue number. With a self-created evaluation article set, PDFDataExtractor achieved promising precision for all key assessed metadata areas of the document text.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Modeling dark- and light-induced crystal structures and single-crystal optical absorption spectra of ruthenium-based complexes that undergo SO 2 -linkage photoisomerization

A family of coordination complexes of the type [Ru(SO 2 )(NH 3 ) 4 X] m+ Y n – (m, n = 1 or 2) exhibit optical switching capabilities in their single-crystal states. This striking effect is caused by the light-induced formation of SO 2 -linkage photoisomers, which are metastable if kept at suitably cool temperatures. We modeled the dark- and light-induced states of these large crystalline complexes via plane-wave (PW)- and molecular-orbital (MO)-based density functional theory (DFT) and time-dependent DFT in order to calculate their structural and optical properties; the calculated results are compared with experimental data. We show that the PW-DFT-based periodic models replicate the structural properties of these complexes more effectively than the MO-DFT-based molecular-fragment models, observing only small deviations in key bond lengths relative to the experimentally derived crystal structures. The periodic models were also found to more effectively simulate trends seen in experimental optical absorption spectra, with optical absorbance and coverage of the visible region increasing with the formation of the photoinduced geometries. The contribution of the metastable photoisomeric species more heavily focuses on the lower-energy end of the spectra. Spectra generated from the molecular-fragment models are limited by the geometry of the fragment used and the number of excited-state roots considered in those calculations. In general, periodic models outperform the molecular-fragment models owing to their ability to better appreciate the periodic phenomena that are present in these crystalline materials as opposed to MO approaches, which are finite methods. We thus demonstrate that PW-DFT-based periodic models should be considered as a more than viable method for simulating the optical and electronic properties of these single-crystal optical switches.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗