Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parsing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Optical emissivity dataset of multi-material heterogeneous designs generated with automated figure extraction

Optical device design is typically an iterative optimization process based on a good initial guess from prior reports. Optical properties databases are useful in this process but difficult to compile because their parsing requires finding relevant papers and manually converting graphical emissivity curves to data tables. Here, we present two contributions: one is a dataset of thermal emissivity records with design-related parameters, and the other is a software tool for automated colored curve data extraction from scientific plots. We manually collected 64 papers with 176 figures reporting thermal emissivity and automatically retrieved 153 colored curve data records. The automated figure analysis software pipeline uses Faster R-CNN for axes and legend object detection, EasyOCR for axes numbering recognition, and k-means clustering for colored curve retrieval. Additionally, we manually extracted geometry, materials, and method information from the text to add necessary metadata to each emissivity curve. Finally, we analyzed the dataset to determine the dominant classes of emissivity curves and determine the underlying design parameters leading to a type of emissivity profile.

47 OTHER INSTRUMENTATION↗

A thermoelectric materials database auto-generated from the scientific literature using ChemDataExtractor

An auto-generated thermoelectric-materials database is presented, containing 22,805 data records, automatically generated from the scientific literature, spanning 10,641 unique extracted chemical names. Each record contains a chemical entity and one of the seminal thermoelectric properties: thermoelectric figure of merit, ZT; thermal conductivity, κ; Seebeck coefficient, S; electrical conductivity, σ; power factor, PF; each linked to their corresponding recorded temperature, T. The database was auto-generated using the automatic sentence-parsing capabilities of the chemistry-aware, natural language processing toolkit, ChemDataExtractor 2.0, adapted for application in the thermoelectric-materials domain, following a rule-based sentence-simplification step. Data were mined from the text of 60,843 scientific papers that were sourced from three scientific publishers: Elsevier, the Royal Society of Chemistry, and Springer. To the best of our knowledge, this is the first automatically-generated database of thermoelectric materials and their properties from existing literature. The database was evaluated to have a precision of 82.25% and has been made publicly available to facilitate the application of data science in the thermoelectric-materials domain, for analysis, design, and prediction.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Decoding substance use disorder severity from clinical notes using a large language model

Substance use disorder (SUD) poses a major concern due to its detrimental effects on health and society. SUD identification and treatment depend on a variety of factors such as severity, co-determinants (e.g., withdrawal symptoms), and social determinants of health. Existing diagnostic coding systems used by insurance providers, like the International Classification of Diseases (ICD-10), lack granularity for certain diagnoses, but American clinicians will add this granularity (as that found within the Diagnostic and Statistical Manual of Mental Disorders classification or DSM-5) as supplemental unstructured text in clinical notes. Traditional natural language processing (NLP) methods face limitations in accurately parsing such diverse clinical language. Large language models (LLMs) offer promise in overcoming these challenges by adapting to diverse language patterns. This study investigates the application of LLMs for extracting severity-related information for various SUD diagnoses from clinical notes. We propose a workflow employing zero-shot learning of LLMs with carefully crafted prompts and post-processing techniques. Through experimentation with Flan-T5, an open-source LLM, we demonstrate its superior recall compared to the rule-based approach. Focusing on 11 categories of SUD diagnoses, we show the effectiveness of LLMs in extracting severity information, contributing to improved risk assessment and treatment planning for SUD patients.

60 APPLIED LIFE SCIENCES↗

Automated electrosynthesis reaction mining with multimodal large language models (MLLMs)

Leveraging the chemical data available in legacy formats such as publications and patents is a significant challenge for the community. Automated reaction mining offers a promising solution to unleash this knowledge into a learnable digital form and therefore help expedite materials and reaction discovery. However, existing reaction mining toolkits are limited to single input modalities (text or images) and cannot effectively integrate heterogeneous data that is scattered across text, tables, and figures. In this work, we go beyond single input modalities and explore multimodal large language models (MLLMs) for the analysis of diverse data inputs for automated electrosynthesis reaction mining. We compiled a test dataset of 65 articles (MERMES-T24 set) and employed it to benchmark five prominent MLLMs against two critical tasks: (i) reaction diagram parsing and (ii) resolving cross-modality data interdependencies. The frontrunner MLLM achieved ≥96% accuracy in both tasks, with the strategic integration of single-shot visual prompts and image pre-processing techniques. We integrate this capability into a toolkit named MERMES (multimodal reaction mining pipeline for electrosynthesis). Our toolkit functions as an end-to-end MLLM-powered pipeline that integrates article retrieval, information extraction and multimodal analysis for streamlining and automating knowledge extraction. This work lays the groundwork for the increased utilization of MLLMs to accelerate the digitization of chemistry knowledge for data-driven research.

Leong, Shi Xuan↗

Candidate strongly lensed type Ia supernovae in the Zwicky Transient Facility archive

Gravitationally lensed type Ia supernovae (glSNe Ia) are unique astronomical tools that can be used to study cosmological parameters, distributions of dark matter, the astrophysics of the supernovae, and the intervening lensing galaxies themselves. A small number of highly magnified glSNe Ia have been discovered by ground-based telescopes such as the Zwicky Transient Facility (ZTF), but simulations predict that a fainter, undetected population may also exist. We present a systematic search for glSNe Ia in the ZTF archive of alerts distributed from June 1 2019 to September 1 2022. Using the AMPEL platform, we developed a pipeline that distinguishes candidate glSNe Ia from other variable sources. Initial cuts were applied to the ZTF alert photometry (with constraints on the peak absolute magnitude and the distance to a catalogue-matched galaxy, as examples) before forced photometry was obtained for the remaining candidates. Additional cuts were applied to refine the candidates based on their light curve colours, lens galaxy colours, and the resulting parameters from fits to the SALT2 SN Ia template. The candidates were also cross-matched with the DESI spectroscopic catalogue. Seven transients were identified that passed all the cuts and had an associated galaxy DESI redshift, which we present as glSN Ia candidates. Although superluminous supernovae (SLSNe) cannot be fully rejected as contaminants, two events, ZTF19abpjicm and ZTF22aahmovu, are significantly different from typical SLSNe and their light curves can be modelled as two-image glSN Ia systems. From this two-image modelling, we estimate time delays of 22 ± 3 and 34 ± 1 days for the two events, respectively, which suggests that we have uncovered a population of glSNe Ia with longer time delays. The pipeline is efficient and sensitive enough to parse full alert streams. It is currently being applied to the live ZTF alert stream to identify and follow-up future candidates while active. This pipeline could be the foundation for glSNe Ia searches in future surveys, such as the Rubin Observatory Legacy Survey of Space and Time.

79 ASTRONOMY AND ASTROPHYSICS↗

WPEC SG50: Developing an Automatically Readable, Comprehensive and Curated Experimental Nuclear Reaction Database

The Organisation for Economic Co-operation and Development (OECD) Nuclear Energy Agency (NEA) Working Party on International Nuclear Data Evaluation Co-operation (WPEC) subgroup (SG) 50 was formed in 2020 to develop an automatically readable, comprehensive and curated experimental nuclear reaction database. This database is called MEDUSAL (Machine-readable Experimental Data User Application & Library), and will draw from EXFOR. The EXFOR database preserves experimental nuclear reaction data true to its original documentation and information from the authors of the data. MEDUSAL will deviate from EXFOR by storing additional information from users of the data for their fields of work (evaluation, model development, validation, etc.). This includes expert judgment on the data sets, identification of data points as outliers, renormalization of the data to the newest monitor reactions, and estimations of missing uncertainty sources. The format for MEDUSAL is being developed to enable easy automatic parsing of large amounts of data. Here, we will summarize the use cases, high-level requirements and first steps towards developing the database MEDUSAL and the API to access it.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

The many-body electronic interactions of Fe(II)–porphyrin

Fe(II)–porphyrin complexes exhibit a diverse range of electronic interactions between the metal and macrocycle. Herein, the incremental full configuration interaction method is applied to the entire space of valence orbitals of a Fe(II)–porphyrin model using a modest basis set. A novel visualization framework is proposed to analyze individual many-body contributions to the correlation energy, providing detailed maps of this complex’s highly correlated electronic structure. Furthermore, this technique is used to parse the numerous interactions of two low-lying triplet states ( 3 A 2 g and 3 E g ) and to show that strong metal d–d and macrocycle π–π orbital interactions preferentially stabilize the 3 A 2 g state. d–π interactions, on the other hand, preferentially stabilize the 3 E g state and primarily appear when correlating six electrons at a time. Ultimately, the Fe(II)–porphyrin model’s full set of 88 valence electrons are correlated in 275 orbitals, showing the interactions up to the 4-body level, which covers the great majority of correlations in this system.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

cclib 2.0: An updated architecture for interoperable computational chemistry

Interoperability in computational chemistry is elusive, impeded by the independent development of software packages and idiosyncratic nature of their output files. The cclib library was introduced in 2006 as an attempt to improve this situation by providing a consistent interface to the results of various quantum chemistry programs. The shared API across programs enabled by cclib has allowed users to focus on results as opposed to output and to combine data from multiple programs or develop generic downstream tools. Initial development, however, did not anticipate the rapid progress of computational capabilities, novel methods, and new programs; nor did it foresee the growing need for customizability. Here, we recount this history and present cclib 2, focused on extensibility and modularity. We also introduce recent design pivots—the formalization of cclib’s intermediate data representation as a tree-based structure, a new combinator-based parser organization, and parsed chemical properties as extensible objects.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Statistical relationships across epigenomes using large-scale hierarchical clustering

Recent advances in genomics and sequencing platforms have revolutionized our ability to create immense data sets, particularly for studying epigenetic regulation of gene expression. However, the avalanche of epigenomic data is difficult to parse for biological interpretation given nonlinear complex patterns and relationships. This attractive challenge in epigenomic data lends itself to machine learning for discerning infectivity and susceptibility. In this study, we explore over 3000 epigenomes of uninfected individuals and provide a framework to characterize the relationships among epigenetic modifiers, their modifiers, genetic loci, and specific immune cell types across all chromosomes using hierarchical clustering. Hierarchical clustering of epigenomic data revealed consistent epigenetic patterns across chromosomes, demonstrating that variation due to epigenetic modifiers is greater than variation between cell types. Gene Ontology and KEGG pathway analyses indicated significant enrichment of genes involved in chromatin remodeling, mRNA splicing, immune responses, and the regulation of microRNAs and snoRNAs. Epigenetic modifiers frequently formed biologically relevant clusters, including the cohesin complex, RNA Polymerase II transcription factors, and PRC2 complex members. These clustering behaviors remained consistent across all chromosomes, supported by entropy analysis and high Adjusted Rand Index scores, indicating robust cross-chromosomal similarity. Co-occurrence analysis further revealed specific sets of modifiers that consistently appeared together within clusters, reflecting shared biological functions and interactions. Validation using another dataset confirmed the reproducibility of these clustering patterns and modifier co-occurrence relationships, underscoring the reliability and generalizability of the methodology.

97 MATHEMATICS AND COMPUTING↗

Meta2DB: Curated Shotgun Metagenomic Feature Sets and Metadata for Health State Prediction

Meta2DB is a curated metagenomic and metadata database that provides structurally consistent microbiome taxonomy feature count tables for 13 897 samples across 84 studies, 23 disease states, and 34 geographical locations. All samples were uniformly processed using a streamlined metagenomic classification pipeline that employs a unique and comprehensive reference database indexed to contain all sequences across all kingdoms of life that were present in the NCBI Nucleotide (nt) database retrieved on 4 January 2023. This pipeline leverages high-performance computing (HPC) resources at Lawrence Livermore National Laboratory and was used to process 50TB of publicly available raw metagenomic sequence data. Extensive metadata curation was carried out through a combination of manual curation and automated parsing, producing a consistent inter-study metadata table specifically structured to facilitate training of ML models for prediction of human health.

Kok, C [Lawrence Livermore National Laboratory (LL↗

A Simple Standard for Sharing Ontological Mappings (SSSOM)

Abstract Despite progress in the development of standards for describing and exchanging scientific information, the lack of easy-to-use standards for mapping between different representations of the same or similar objects in different databases poses a major impediment to data integration and interoperability. Mappings often lack the metadata needed to be correctly interpreted and applied. For example, are two terms equivalent or merely related? Are they narrow or broad matches? Or are they associated in some other way? Such relationships between the mapped terms are often not documented, which leads to incorrect assumptions and makes them hard to use in scenarios that require a high degree of precision (such as diagnostics or risk prediction). Furthermore, the lack of descriptions of how mappings were done makes it hard to combine and reconcile mappings, particularly curated and automated ones. We have developed the Simple Standard for Sharing Ontological Mappings (SSSOM) which addresses these problems by: (i) Introducing a machine-readable and extensible vocabulary to describe metadata that makes imprecision, inaccuracy and incompleteness in mappings explicit. (ii) Defining an easy-to-use simple table-based format that can be integrated into existing data science pipelines without the need to parse or query ontologies, and that integrates seamlessly with Linked Data principles. (iii) Implementing open and community-driven collaborative workflows that are designed to evolve the standard continuously to address changing requirements and mapping practices. (iv) Providing reference tools and software libraries for working with the standard. In this paper, we present the SSSOM standard, describe several use cases in detail and survey some of the existing work on standardizing the exchange of mappings, with the goal of making mappings Findable, Accessible, Interoperable and Reusable (FAIR). The SSSOM specification can be found at http://w3id.org/sssom/spec. Database URL: http://w3id.org/sssom/spec

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

pyXPCSviewer : an open-source interactive tool for X-ray photon correlation spectroscopy visualization and analysis

pyXPCSviewer , a Python-based graphical user interface that is deployed at beamline 8-ID-I of the Advanced Photon Source for interactive visualization of XPCS results, is introduced. pyXPCSviewer parses rich X-ray photon correlation spectroscopy (XPCS) results into independent PyQt widgets that are both interactive and easy to maintain. pyXPCSviewer is open-source and is open to customization by the XPCS community for ingestion of diversified data structures and inclusion of novel XPCS techniques, both of which are growing demands particularly with the dawn of near-diffraction-limited synchrotron sources and their dedicated XPCS beamlines.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

To Derive or Not to Derive: I/O Libraries Take Charge of Derived Quantities Computation

The ever-increasing volume of data produced by HPC simulations necessitates scalable methods for data exploration and knowledge extraction. Scientific data analysis often involves complex queries across distributed datasets, requiring manipulation of multiple primary variables and generating derived data that needs to be handled efficiently, creating challenges for applications that need to parse many large datasets. Relying on individual applications to handle all intermediate data generally leads to redundant computations across studies and unnecessary data transfers. In this paper, we investigate the performance of different approaches where applications define derived variables as quantities of interest (QoIs) and offload the computation and transfer of these QoIs to the I/O library. This significantly reduces redundancy and optimizes data movement across the distributed storage and processing infrastructure by allowing control over when and where derived variables are computed. We present a detailed analysis of the performance-storage trade-offs associated with different solutions and showcase results for our study on two large-scale datasets created from climate and combustion simulations.

Gainaru, Ana↗

Development of Real-Time High-Density Pulsar Data Transmission and Processing for Grid Synchronization

Taking advantage of the extreme stability of the pulsar period, it can serve as the timing source for grid synchronization to compensate for the timing drift instigated by the loss of GPS signal. Nevertheless, the real-time transmission and processing of the pulsar data suffer from its high-frequency data rate, varying from megahertz to gigahertz, resulting in reduced computing speed and increased time delay. To mitigate this issue, the hardware and software frameworks are implemented for the high-density pulsar data transmission and processing for grid synchronization in this research. Initially, the high-density pulsar data is transferred using open-source software. The complementary duty cycle timing module is designed to coordinate the operation of the dual-channel high-speed interface and software. Subsequently, the multiple-threading is applied to the receiving, parsing, and splicing pulsar data. Next, the pulsar signal extraction method is implemented based on the polyphase filterbank and time of arrival estimation. Ultimately, real-time performance verification experiments are carried out for different components under two hardware platforms. Finally, the results demonstrate that only 0.482 s is required for processing 4 Gigabyte data through multiple-threading, which is 3.8 times faster than the single thread. The pulsar signal extraction can also be executed within 707 ms for 4.8 seconds of data, thereby indicating that real-time requirements can be met.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Sparse Symmetric Format for Tucker Decomposition

Tensor-based methods are receiving renewed attention in recent years due to their prevalence in diverse real-world applications. There is considerable literature on tensor representations and algorithms for tensor decompositions, both for dense and sparse tensors. Many applications in hypergraph analytics, machine learning, psychometry, and signal processing result in tensors that are both sparse and symmetric, making them an important class for further study. Similar to the critical Tensor Times Matrix chain operation (TTM c ) in general sparse tensors, the $\underline{S}$ parse $\underline{S}$ ymmetric $\underline{T}$ ensor $\underline{T}$ imes $\underline{S}$ ame $\underline{M}$ atrix $\underline{c}$ hain (S 3 TTM c ) operation is compute and memory intensive due to high tensor order and the associated factorial explosion in the number of non-zeros. We present the novel Compressed Sparse Symmetric (CSS) format for sparse symmetric tensors, along with an efficient parallel algorithm for the S 3 TTM c operation. We theoretically establish that S 3 TTM c on CSS achieves a better memory versus run-time trade-off compared to state-of-the-art implementations, and visualize the variation of the performance gap over the parameter space. We demonstrate experimental findings that confirm these results and achieve up to 2.72× speedup on synthetic and real datasets. The scaling of the algorithm on different test architectures is also showcased to highlight the effect of machine characteristics on algorithm performance.

42 ENGINEERING↗

Interactions Between Climate Mean and Variability Drive Future Agroecosystem Vulnerability

ABSTRACT Agriculture is crucial for global food supply and dominates the Earth's land surface. It is unknown, however, how slow but relentless changes in climate mean state, versus random extreme conditions arising from changing variability , will affect agroecosystems' carbon fluxes, energy fluxes, and crop production. We used an advanced weather generator to partition changes in mean climate state versus variability for both temperature and precipitation, producing forcing data to drive factorial‐design simulations of US Midwest agricultural regions in the Energy Exascale Earth System Model. We found that an increase in temperature mean lowers stored carbon, plant productivity, and crop yield, and tends to convert agroecosystems from a carbon sink to a source, as expected; it also can cause local to regional cooling in the earth system model through its effects on the Bowen Ratio. The combined effect of mean and variability changes on carbon fluxes and pools was nonlinear, that is, greater than each individual case. For instance, gross primary production reduces by 9%, 1%, and 13% due to change in mean temperature, change in temperature variability, and change in both temperature mean and variability, respectively. Overall, the scenario with change in both temperature and precipitation means leads to the largest reduction in carbon fluxes (−16% gross primary production), carbon pools (−35% vegetation carbon), and crop yields (−33% and −22% median reduction in yield for corn and soybean, respectively). By unambiguously parsing the effects of changing climate mean versus variability and quantifying their nonadditive impacts, this study lays a foundation for more robust understanding and prediction of agroecosystems' vulnerability to 21st‐century climate change.

54 ENVIRONMENTAL SCIENCES↗

Variation in Root Exudate Composition Influences Soil Microbiome Membership and Function

Root exudation is one of the primary processes that mediate interactions between plant roots, microorganisms, and the soil matrix, yet the mechanisms by which exudation alters microbial metabolism in soils have been challenging to unravel. Here, utilizing distinct sorghum genotypes, we characterized the chemical heterogeneity between root exudates and the effects of that variability on soil microbial membership and metabolism. Distinct exudate chemical profiles were quantified and used to formulate synthetic root exudate treatments: a high-organic-acid treatment (HOT) and a high-sugar treatment (HST). To parse the response of the soil microbiome to different exudate regimens, laboratory soil reactors were amended with these root exudate treatments as well as a nonexudate control. Amplicon sequencing of the 16S rRNA gene illustrated distinct microbial diversity patterns and membership in response to HST, HOT, or control amendments. Exometabolite changes reflected these microbial community changes, and we observed enrichment of organic and amino acids, as well as possible phytohormones in the HST relative to the HOT and control. Linking the metabolic capacity of metagenome-assembled genomes in the HST to the exometabolite patterns, we identified microorganisms that could produce these phytohormones. Our findings emphasize the tractability of high-resolution multiomics tools to investigate soil microbiomes, opening the possibility of manipulating native microbial communities to improve specific soil microbial functions and enhance crop production.

59 BASIC BIOLOGICAL SCIENCES↗

New insights into the natural history of bronchopulmonary dysplasia from proteomics and multiplexed immunohistochemistry

Bronchopulmonary dysplasia (BPD) is a disease of prematurity related to the arrest of normal lung development. The objective of this study was to better understand how proteome modulation and cell-type shifts are noted in BPD pathology. Pediatric human donors aged 1–3 yr were classified based on history of prematurity and histopathology consistent with “healed” BPD (hBPD, n = 3) and “established” BPD (eBPD, n = 3) compared with respective full-term born (n = 6) age-matched term controls. Proteins were quantified by tandem mass spectroscopy with selected Western blot validations. Multiplexed immunofluorescence (MxIF) microscopy was performed on lung sections to enumerate cell types. Protein abundances and MxIF cell frequencies were compared among groups using ANOVA. Cell type and ontology enrichment were performed using an in-house tool and/or EnrichR. Proteomics detected 5,746 unique proteins, 186 upregulated and 534 downregulated, in eBPD versus control with fewer proteins differentially abundant in hBPD as compared with age-matched term controls. Cell-type enrichment suggested a loss of alveolar type I, alveolar type II, endothelial/capillary, and lymphatics, and an increase in smooth muscle and fibroblasts consistent with MxIF. Histochemistry and Western analysis also supported predictions of upregulated ferroptosis in eBPD versus control. Finally, several extracellular matrix components mapping to angiogenesis signaling pathways were altered in eBPD. Despite clear parsing by protein abundance, comparative MxIF analysis confirms phenotypic variability in BPD. This work provides the first demonstration of tandem mass spectrometry and multiplexed molecular analysis of human lung tissue for critical elucidation of BPD trajectory-defining factors into early childhood.

60 APPLIED LIFE SCIENCES↗