Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “vocabulary”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Path-BigBird: An AI-Driven Transformer Approach to Classification of Cancer Pathology Reports

PURPOSE Surgical pathology reports are critical for cancer diagnosis and management. To accurately extract information about tumor characteristics from pathology reports in near real time, we explore the impact of using domain-specific transformer models that understand cancer pathology reports. METHODS We built a pathology transformer model, Path-BigBird, by using 2.7 million pathology reports from six SEER cancer registries. We then compare different variations of Path-BigBird with two less computationally intensive methods: Hierarchical Self-Attention Network (HiSAN) classification model and an offthe-shelf clinical transformer model (Clinical BigBird). We use five pathology information extraction tasks for evaluation: site, subsite, laterality, histology, and behavior. Model performance is evaluated by using macro and micro F 1 scores. RESULTS We found that Path-BigBird and Clinical BigBird outperformed the HiSAN in all tasks. Clinical BigBird performed better on the site and laterality tasks. Versions of the Path-BigBird model performed best on the two most difficult tasks: subsite (micro F 1 score of 72.53, macro F 1 score of 35.76) and histology (micro F 1 score of 80.96, macro F 1 score of 37.94). The largest performance gains over the HiSAN model were for histology, for which a Path-BigBird model increased the micro F 1 score by 1.44 points and the macro F 1 score by 3.55 points. Overall, the results suggest that a Path-BigBird model with a vocabulary derived from wellcurated and deidentified data is the best-performing model. CONCLUSION The Path-BigBird pathology transformer model improves automated information extraction from pathology reports. Although Path-BigBird outperforms Clinical BigBird and HiSAN, these less computationally expensive models still have utility when resources are constrained.

60 APPLIED LIFE SCIENCES↗

Tokenized Data for FORGE Foundation Models

This dataset comprises a vast corpus of 257 billion tokens, accompanied by the corresponding vocabulary file employed in the pre-training of FORGE foundation models. The primary data source for this corpus is scientific documents derived from diverse origins, and they have been tokenized using the Hugging Face BPE tokenizer. Further details about this research can be found in the publication titled FORGE: Pre-Training Open Foundation Models for Science authored by Junqi Yin, Sajal Dash, Feiyi Wang, and Mallikarjun (Arjun) Shankar, presented at SC'23. The data tokenization pipeline and resulting artifacts use CORE data [Ref: Knoth, P., and Zdrahal, Z. (2012). CORE: three access levels to underpin open access. D-Lib Magazine, 18(11/12)]. For use of these data sets for any purpose, please follow the guidelines provided in https://core.ac.uk/terms .

Yin, Junqi↗

FAIRLinked: Data FAIRification Tools for Materials Data Science

FAIRLinked is a software package created to support the FAIRification of materials science data, ensuring proper alignment with FAIR principles: Findable, Accessible, Interoperable, and Reusable. It is built to be compatible with MDS-Onto, an ontology designed to capture the semantics of various types of materials data, enabling integration and sharing across different research workflows. The package is subdivided into three subpackages: InterfaceMDS, RDFTableConversion, and QBWorkflow. The first subpackage, InterfaceMDS allows users to search for terms using either string search or various filters, explore different domains and subdomains, and add terms to MDS-Onto. RDFTableConversion is used for serialization and deserialization of data from CSV into JSONLDs and vice versa in a way that captures the semantics of the data using MDS-Onto. Lastly, QBWorkflow is a serialization and deserialization workflow that incorporates RDF Data Cube vocabulary, useful for working with multidimensional datasets. By offering these packages, FAIRLinked lowers the barrier of creating FAIR, machine-actionable data for researchers in the materials science community.

FAIR↗

A Power Application Developer’s Guide to the Common Information Model: An Introduction for Power Systems Engineers and Application Developers – CIM17v40

A key issue in creating the next generation of energy management system (EMS) and advanced distribution management system (ADMS) platforms will be the ability to represent and exchange power system network model data in a consistent manner. To this end, the Common Information Model (CIM) stands out as the only standardized vocabulary (or ontology) for defining power system network models and asset data in a comprehensive, consistent manner across the generation-transmission-distribution boundary. The CIM is freely available to use and extend. The CIM is maintained by the UCAiug (informally known as the CIM User’s Group) under an Apache 2.0 license. The CIM Users Group collaborates with the IEC and other standards communities for the development of technical and informative specifications. Although portions of the information model are referred to by the corresponding IEC standards naming, it is not necessary to purchase any of the IEC standards to use the CIM information model. This document provides a roadmap for power system engineers and application developers not familiar with semantic modeling to start using the CIM for modeling, simulation, optimization, and development of advanced power applications. The key classes needed for defining power system topology and equipment are explained systematically. Key focus areas include modeling of lines, transformers, generators, switching equipment, loads, and distributed energy resources (DERs).

97 MATHEMATICS AND COMPUTING↗

Skewering the silos: using Brick to enable portable analytics, modeling and controls in buildings

Nearly all large commercial buildings have heating, ventilation and air conditioning (HVAC) systems, lighting systems, safety and other systems controlled by a computer—a dedicated server with a building energy management system (BMS). However, these BMSs are proprietary with each building’s assets (that is, fans, valves, pumps, and their setpoints) named and coded uniquely by the BMS vendor or engineer; building analytics and control algorithms are written specific to the assets and the building. Thus, any control updates or analytics to improve building performance—especially critical to reduce greenhouse emissions or improve load flexibility—are labor intensive and costly. The Brick schema was developed so the same analysis or control algorithms can work on a variety of buildings if each is digitally represented in a Brick data model. The goal of this project was to further the development of Brick to extend it beyond an academic project with demonstrated success in a small field study, to a practical choice for industrial and commercial stakeholders seeking to realize value from building data. To do this, we executed four objectives: (1) expand the Brick schema including its modeling capabilities and vocabulary, (2) develop tools for integrating Brick with existing digital technologies and representations in buildings, (3) develop an open-source analytics platform to facilitate use of Brick in delivering data value, and (4) demonstrate Brick-driven analytics and controls in real settings.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Improving the Nomenclature Around Uncertainty [Slides]

This presentation finds that defining the marginal probability density function (PDF) for nuclear data is important. Additionally, the vocabulary of “means and covariances” and new GNDS 2.0 formats are limited to Gaussian (normal) representations— always incorrect—but clearly of practical significance when uncertainties are large (>40%). Finally, the Triage Solution: declare our current data as containing best estimate (mode) plus variance for a truncated normal or lognormal.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Metadata Standards for the NSE: Core Fields

This standard presents a core set of metadata fields required for each managed digital object within the Nuclear Security Enterprise (NSE). Metadata standardization is a critical enabler for two primary objectives: 1) effectively sharing data, documents, and other digital objects between NSE sites; and 2) supporting digital engineering through the digital thread at the object level. Standardization is necessary for both schemas and vocabularies, meaning that both field standards and value standards must be specified. This document serves as a foundational field standard, recommending a core set of fields that should be uniformly required for all managed digital objects within the NSE.

99 GENERAL AND MISCELLANEOUS↗

HPC-FAIR: A Framework Managing Data and AI Models for Analyzing and Optimizing Scientific Applications

The increasing reliance on machine learning (ML) to analyze and optimize large-scale scientific applications on supercomputers faces a significant bottleneck: the lack of readily available, high-quality training datasets and the difficulty in reusing existing AI models. This project was motivated by the urgent need to address the “FAIR” principles (Findability, Accessibility, Interoperability, Reusability) for both training datasets and AI models in the high-performance computing (HPC) domain. The project developed HPC-FAIR, a high-performance computing data management framework designed to centralize HPC-related datasets and AI models within a unified hub. To ensure interoperability, the framework established a standardized representation and vocabulary (ontology) for both data and models. HPC-FAIR also implemented automated workflows to streamline data processing, model access, and benchmarking. Additionally, the project focused on optimizing data harnessing efficiency through advanced techniques like deep reuse and compression-based analytics.

97 MATHEMATICS AND COMPUTING↗

NEPATEC v2.0: Standardized Metadata and Text Corpus of National Environmental Policy Act Documents

The National Environmental Policy Act of 1969, as amended (NEPA), is a major environmental law in the United States, requiring Federal agencies to consider and document potential environmental impacts before deciding on a proposed action. Modernization of NEPA and permitting processes faces significant challenges due to the lack of standardized formats and interoperable systems for organizing and sharing NEPA-related information across agencies. Much of the information gathered during NEPA reviews is written into documents such as categorical exclusions, environmental assessments, and environmental impact statements, then filed in predominately independent agency file stores that may or may not be publicly accessible. The application of metadata and data standards, such as those recommended by the Council on Environmental Quality (CEQ), to NEPA documents offers a shared vocabulary and structure for key entities like projects, processes, and documents that can streamline information exchange and enhance collaboration across systems. In this work, we publicly release NEPATEC2.0, an expanded corpus of NEPA documents with associated metadata. NEPATEC2.0 encompasses approximately 120,000 documents from 60,000 projects prepared by more than 60 different agencies. Modeled to align with CEQ metadata standards, NEPATEC2.0 promotes consistency in environmental reviews and supports the ongoing effort to modernize permitting technologies by facilitating more transparent, efficient, and data-driven decision-making. Importantly, NEPATEC2.0 demonstrates the possibilities and limitations of large language model-based prompting to extract information from NEPA documents at scale.

54 ENVIRONMENTAL SCIENCES↗

BASIN-3D Data Integration for Selected ARM Data Field Campaign Report

The purpose of this data services request was to demonstrate integration of the Atmospheric Radiation Measurement (ARM) User Facility’s “met” datastreams with time series data from other earth science data sources using the BASIN-3D data synthesis software tool. BASIN-3D is an open-source Python library that enables researchers to integrate data across configured public and private data sources. It provides a common query language for researchers to request measurement locations and time series data based on specified locations, variables, time period, statistics, aggregation, and data quality. BASIN-3D acquires the data that match the query from each configured data source and translates the results into harmonized vocabularies, thus reducing researchers' data-wrangling effort. In addition, because the queries are executed on demand, researchers can easily regenerate their synthesized data sets as new data and/or data updates become available, eliminating one-off data products. BASIN-3D can output data using a variety of different data structures for end-user applications including Python pandas data frames and hdf5 output formats.

54 ENVIRONMENTAL SCIENCES↗

Models and Processes to Extract Drug-like Molecules From Natural Language Text

Researchers worldwide are seeking to repurpose existing drugs or discover new drugs to counter the disease caused by severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2). A promising source of candidates for such studies is molecules that have been reported in the scientific literature to be drug-like in the context of viral research. However, this literature is too large for human review and features unusual vocabularies for which existing named entity recognition (NER) models are ineffective. We report here on a project that leverages both human and artificial intelligence to detect references to such molecules in free text. We present 1) a iterative model-in-the-loop method that makes judicious use of scarce human expertise in generating training data for a NER model, and 2) the application and evaluation of this method to the problem of identifying drug-like molecules in the COVID-19 Open Research Dataset Challenge (CORD-19) corpus of 198,875 papers. We show that by repeatedly presenting human labelers only with samples for which an evolving NER model is uncertain, our human-machine hybrid pipeline requires only modest amounts of non-expert human labeling time (tens of hours to label 1778 samples) to generate an NER model with an F-1 score of 80.5%—on par with that of non-expert humans—and when applied to CORD’19, identifies 10,912 putative drug-like molecules. This enriched the computational screening team’s targets by 3,591 molecules, of which 18 ranked in the top 0.1% of all 6.6 million molecules screened for docking against the 3CLPro protein.

60 APPLIED LIFE SCIENCES↗

The relationship between below average cognitive ability at age 5 years and the child’s experience of school at age 9

Background At age 5, while only embarking on their educational journey, substantial differences in children’s cognitive ability will already exist. The aim of this study was to examine the causal association between below average cognitive ability at age 5 years and child-reported experience of school and self-concept, and teacher-reported class engagement and emotional-behavioural function at age 9 years. Methods This longitudinal cohort study used data from 7,392 children in the Growing Up in Ireland Infant Cohort, who had completed the Picture Similarities and Naming Vocabulary subtests of the British Abilities Scales at age 5. Principal components analysis was used to produce a composite general cognitive ability score for each child. Children with a general cognitive ability score more than 1 standard deviation (SD) below the mean at age 5 were categorised as ‘Below Average Cognitive Ability’ (BACA), and those scoring above this as ‘Typical Cognitive Development’ (TCD). The outcomes of interest, measured at age 9, were child-reported experience of school, child’s self-concept, teacher-reported class engagement, and teacher-reported emotional behavioural function. Binary and multinomial logistic regression models were used to examine the association between BACA and these outcomes. Results Compared to those with TCD, those with BACA had significantly higher odds of never liking school [Adjusted odds ratio (AOR) 1.82, 95% CI 1.37–2.43, p < 0.001], of being picked on (AOR 1.27, 95% CI 1.09–1.48) and of picking on others (AOR 1.53, 95% CI 1.27–1.84). They had significantly higher odds of experiencing low self-concept (AOR 1.20, 95% CI 1.02–1.42) and emotional-behavioural difficulties (AOR 1.34, 95% CI 1.10–1.63, p = 0.003). Compared to those with TCD, children with BACA had significantly higher odds of hardly ever or never being interested, motivated and excited to learn (AOR 2.29, 95% CI 1.70–3.10). Conclusion Children with BACA at school-entry had significantly higher odds of reporting a negative school experience and low self-concept at age 9. They had significantly higher odds of having teacher-reported poor class engagement and problematic emotional-behavioural function at age 9. The findings of this study suggest BACA has a causal role in these adverse outcomes. Early childhood policy and intervention design should be cognisant of the important role of cognitive ability in school and childhood outcomes.

Bowe, Andrea K.↗

Development of a Unified Taxonomy for HVAC System Faults

Detecting and diagnosing HVAC faults is critical for maintaining building operation performance, reducing energy waste, and ensuring indoor comfort. An increasing deployment of commercial fault detection and diagnostics (FDD) software tools in commercial buildings in the past decade has significantly increased buildings’ operational reliability and reduced energy consumption. A massive amount of data has been generated by the FDD software tools. However, efficiently utilizing FDD data for ‘big data’ analytics, algorithm improvement, and other data-driven applications is challenging because the format and naming conventions of those data are very customized, unstructured, and hard to interpret. This paper presents the development of a unified taxonomy for HVAC faults. A taxonomy is an orderly classification of HVAC faults according to their characteristics and causal relations. The taxonomy includes fault categorization, physical hierarchy, fault library, relation model, and naming/tagging scheme. The taxonomy employs both a physical hierarchy of HVAC equipment and a cause-effect relationship model to reveal the root causes of faults in HVAC systems. A structured and standardized vocabulary library is developed to increase data representability and interpretability. The developed fault taxonomy can be used for HVAC system ‘big data’ analytics such as HVAC system fault prevalence analysis or the development of an HVAC FDD software standard. A common type of HVAC equipment-packaged rooftop unit (RTU) is used as an example to demonstrate the application of the developed fault taxonomy. Two RTU FDD software tools are used to show that after mapping FDD data according to the taxonomy, the meta-analysis of the multiple FDD reports is possible and efficient.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Reviews and syntheses: The promise of big diverse soil data, moving current practices towards future potential

Abstract. In the age of big data, soil data are more available and richer than ever, but – outside of a few large soil survey resources – they remain largely unusable for informing soil management and understanding Earth system processes beyond the original study. Data science has promised a fully reusable research pipeline where data from past studies are used to contextualize new findings and reanalyzed for new insight. Yet synthesis projects encounter challenges at all steps of the data reuse pipeline, including unavailable data, labor-intensive transcription of datasets, incomplete metadata, and a lack of communication between collaborators. Here, using insights from a diversity of soil, data, and climate scientists, we summarize current practices in soil data synthesis across all stages of database creation: availability, input, harmonization, curation, and publication. We then suggest new soil-focused semantic tools to improve existing data pipelines, such as ontologies, vocabulary lists, and community practices. Our goal is to provide the soil data community with an overview of current practices in soil data and where we need to go to fully leverage big data to solve soil problems in the next century.

54 ENVIRONMENTAL SCIENCES↗

Ground Operations Aerospace Language (GOAL). Volume 5: Application Studies

The Ground Operations Aerospace Language (GOAL) was designed to be used by test oriented personnel to write procedures which would be executed in a test environment. A series of discussions between NASA LV-CAP personnel and IBM resulted in some peripheral tasks which would aid in evaluating the applicability of the language in this environment, and provide enhancement for future applications. The results of these tasks are contained within this volume. The GOAL vocabulary provides a high degree of readability and retainability. To achieve these benefits, however, the procedure writer utilizes words and phrases of considerable length. Brief form study was undertaken to determine a means of relieving this burden. The study resulted in a version of GOAL which enables the writer to develop a dialect suitable to his needs and satisfy the syntax equations. The output of the compiler would continue to provide readability by printing out the standard GOAL language. This task is described.

Source record↗

Crew/computer communications study. Volume 1: Final report

Techniques, methods, and system requirements are reported for an onboard computerized communications system that provides on-line computing capability during manned space exploration. Communications between man and computer take place by sequential execution of each discrete step of a procedure, by interactive progression through a tree-type structure to initiate tasks or by interactive optimization of a task requiring man to furnish a set of parameters. Effective communication between astronaut and computer utilizes structured vocabulary techniques and a word recognition system.

Johannes, J. D.↗