Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “vocabulary”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

NASA's Global Change Master Directory: Discover and Access Earth Science Data Sets, Related Data Services, and Climate Diagnostics

NASA's Global Change Master Directory provides the scientific community with the ability to discover, access, and use Earth science data, data-related services, and climate diagnostics worldwide. The GCMD offers descriptions of Earth science data sets using the Directory Interchange Format (DIF) metadata standard; Earth science related data services are described using the Service Entry Resource Format (SERF); and climate visualizations are described using the Climate Diagnostic (CD) standard. The DIF, SERF and CD standards each capture data attributes used to determine whether a data set, service, or climate visualization is relevant to a user's needs. Metadata fields include: title, summary, science keywords, service keywords, data center, data set citation, personnel, instrument, platform, quality, related URL, temporal and spatial coverage, data resolution and distribution information. In addition, nine valuable sets of controlled vocabularies have been developed to assist users in normalizing the search for data descriptions. An update to the GCMD's search functionality is planned to further capitalize on the controlled vocabularies during database queries. By implementing a dynamic keyword "tree", users will have the ability to search for data sets by combining keywords in new ways. This will allow users to conduct more relevant and efficient database searches to support the free exchange and re-use of Earth science data. http://gcmd.nasa.gov/

Aleman, Alicia↗

A Survey of Methods for Computing Best Estimates of Endoatmospheric and Exoatmospheric Trajectories

Beginning with the mathematical prediction of planetary orbits in the early seventeenth century up through the most recent developments in sensor fusion methods, many techniques have emerged that can be employed on the problem of endo and exoatmospheric trajectory estimation. Although early methods were ad hoc, the twentieth century saw the emergence of many systematic approaches to estimation theory that produced a wealth of useful techniques. The broad genesis of estimation theory has resulted in an equally broad array of mathematical principles, methods and vocabulary. Among the fundamental ideas and methods that are briefly touched on are batch and sequential processing, smoothing, estimation, and prediction, sensor fusion, sensor fusion architectures, data association, Bayesian and non Bayesian filtering, the family of Kalman filters, models of the dynamics of the phases of a rocket's flight, and asynchronous, delayed, and asequent data. Along the way, a few trajectory estimation issues are addressed and much of the vocabulary is defined.

Bernard, William P.↗

Reports of Resilient Performance: Investigating Operators' Descriptions of Safety-producing Behaviors in the Aviation Safety Reporting System

While many existing taxonomies and frameworks provide a common vocabulary for describing how human operators fail in the context of sociotechnical systems, at present, there is no common vocabulary to describe how humans succeed. Such a framework would facilitate systematically collecting and analyzing data on how human performance can produce safety, not just how it can reduce safety. One potentially rich source of currently available information for exploring desired performance is the reports submitted to NASA’s Aviation Safety Reporting System (ASRS). These de-identified, confidential, and voluntary narrative reports are submitted by pilots, controllers, ground operators, and others within aviation operations. While these reports are primarily submitted to describe safety risks, incidents, and problems, they also often describe how those risks were mitigated, and provide a window into aspects of everyday work in aviation. This paper describes an analysis of ASRS narratives to understand how operators talk about their own resilient behaviors during adverse safety conditions and events. Guided by Erik Hollnagel’s Resilience Assessment Grid framework (i.e., anticipate, monitor, respond, learn), we illustrate our approach and methodology with examples from reports. We also highlight some of the challenges and how further research is needed in developing a taxonomy of operators’ descriptions of resilient performance.

resilient behaviors↗

Reports of Resilient Performance: Investigating Operators' Descriptions of Safety-producing Behaviors in the Aviation Safety Reporting System

While many existing taxonomies and frameworks provide a common vocabulary for describing how human operators fail in the context of sociotechnical systems, at present, there is no common vocabulary to describe how humans succeed. Such a framework would facilitate systematically collecting and analyzing data on how human performance can produce safety, not just how it can reduce safety. One potentially rich source of currently available information for exploring desired performance is the reports submitted to NASA’s Aviation Safety Reporting System (ASRS). These de-identified, confidential, and voluntary narrative reports are submitted by pilots, controllers, ground operators, and others within aviation operations. While these reports are primarily submitted to describe safety risks, incidents, and problems, they also often describe how those risks were mitigated, and provide a window into aspects of everyday work in aviation. These reports can be searched in a variety of ways. This paper describes methods for systematically examining ASRS narratives to understand how operators talk about their own resilient behaviors during adverse safety conditions and events. Guided by Erik Hollnagel’s Resilience Assessment Grid framework (i.e., anticipate, monitor, respond, learn), various approaches and tools for such inquiries are described. The approach, process, challenges, and suggestions to building an operator-based description of resilient performance are discussed, and tools to facilitate data analysis to maximize learning from these reports are described.

ASRS narratives↗

Keywords for All People: How Keyword Governance and Coordination with the NASA ESDIS Standards Coordination Office (ESCO) Improves GCMD Keywords for Discuvery and Use

The Global Change Master Directory (GCMD) Keywords, initiated over twenty years ago, are a hierarchical set of controlled Earth Science vocabularies that help ensure Earth science data, services, and variables are described in a consistent and comprehensive manner and allow for the precise searching of metadata and subsequent retrieval of data, services, and variables. GCMD keywords are periodically analyzed for relevancy and will continue to be refined and expanded in response to user needs. The periodic analysis is a result of successful coordination with the ESDIS Standards Coordination Office (ESCO), which is responsible for standards activities across ESDIS, and assists with providing valuable stakeholder and subject matter expert (SME) feedback on GCMD vocabularies. The ESCO is also turning to the GCMD to discuss how the keyword review process can be improved and more streamlined. In addition to the ESCO, keyword requests and feedback are also received through the GCMD Keyword Forum.

Tyler Stevens↗

Metadata Standards for the NSE: Extended Field Standards

This standard presents a set of optional metadata fields for managed digital objects within the Nuclear Security Enterprise (NSE) and provides a deeper look at data representation in metadata by looking at the representation of 1) Records Management required metadata, and 2) common representations of technical/scientific data. Metadata standardization is a critical enabler for effectively sharing data, documents, and other digital objects between NSE sites, and for tracing the digital thread at the object level. Standardization is necessary for both schemas and vocabularies, meaning that both field standards and value standards must be specified. This document serves as a complementary field standard, recommending an optional set of fields that should be uniformly built for all managed digital objects within the NSE. This document specifically focuses on extending the shared discovery layer defined in the first white paper by introducing additional descriptive and data representation fields that improve cross-site search and interpretation.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Global Change Master Directory (GCMD) Keyword Management Process and Lifecycle

The Global Change Master Directory (GCMD) keywords are a hierarchical set of controlled vocabulary covering the Earth science disciplines that have been evolving for over 25 years. The process for how these keywords have been curated and reviewed has also evolved. This presentation will convey the process for reviewing and approving the GCMD keywords, including fast track and yearly reviews through the ESDIS Standards Office, and how Earth science users can influence keyword additions and modifications. The presentation will also highlight how keywords facilitate the discovery of EOSDIS data and services and how other organizations are using the GCMD keywords.

EOSDIS↗

Explainable machine learning reveals that local structural motifs encode the thermodynamic state across the CuZr metallic glass-forming range

Metallic glasses derive their properties from the statistics of local atomic motifs rather than from long-range order, yet a quantitative, chemistry-specific link between motif populations and the underlying glassy state has remained elusive. In this work we combine large-scale molecular dynamics, Voronoi tessellation, deep neural networks, and SHapley Additive exPlanations (SHAP) to identify which local structural motifs define the glassy state of Cu—Zr metallic glasses. A dataset of 17,180 atomistic configurations spanning ten compositions (Cu 20 Zr 80 –Cu 80 Zr 20 ) and four quench rates (10 9 –10 12 K/s) is used to train a feed-forward neural network that regresses temperature across the 50–2000 K liquid–supercooled–glass range, achieving a mean absolute error of 19.89 K and R 2 = 0.9974, confirming that the local structural state is faithfully encoded in motif-level structure. SHAP analysis then reveals that a tightly coupled near-icosahedral family of motifs (coordination numbers (CN) 11–13, including the full icosahedron 001200 and its single-atom-perturbation sibling 10930) collectively encodes the thermodynamic state of the system across the full glass-forming range. The CN = 11–13 ordered members carry negative SHAP values at high populations, tracking the most deeply-quenched configurations, while 10930 shows the reversed signature consistent with its role as a soft-spot host whose population shrinks as the icosahedral network deepens. The analysis demonstrates that explainable machine learning can isolate the minimal motif vocabulary defining the glassy state and recovers the near-icosahedral building blocks previously identified by data-driven analyses of Cu—Zr. The approach provides a general, chemistry-specific route for characterizing the structural state of disordered materials.

36 MATERIALS SCIENCE↗

Characterizing and communicating uncertainty: lessons from NASA’s Carbon Monitoring System

Navigating uncertainty is a critical challenge in all fields of science, especially when translating knowledge into real-world policies or management decisions. However, the wide variance in concepts and definitions of uncertainty across scientific fields hinders effective communication. As a microcosm of diverse fields within Earth Science, NASA’s Carbon Monitoring System (CMS) provides a useful crucible in which to identify cross-cutting concepts of uncertainty. The CMS convened the Uncertainty Working Group (UWG), a group of specialists across disciplines, to evaluate and synthesize efforts to characterize uncertainty in CMS projects. This paper represents efforts by the UWG to build a heuristic framework designed to evaluate data products and communicate uncertainty to both scientific and non-scientific end users. We consider four pillars of uncertainty: origins, severity, stochasticity versus incomplete knowledge, and spatial and temporal autocorrelation. Using a common vocabulary and a generalized workflow, the framework introduces a graphical heuristic accompanied by a narrative, exemplified through contrasting case studies. Envisioned as a versatile tool, this framework provides clarity in reporting uncertainty, guiding users and tempering expectations. Beyond CMS, it stands as a simple yet powerful means to communicate uncertainty across diverse scientific communities.

54 ENVIRONMENTAL SCIENCES↗

Structured Prompt Interrogation and Recursive Extraction of Semantics (SPIRES): a method for populating knowledge bases using zero-shot learning

Abstract Motivation Creating knowledge bases and ontologies is a time consuming task that relies on manual curation. AI/NLP approaches can assist expert curators in populating these knowledge bases, but current approaches rely on extensive training data, and are not able to populate arbitrarily complex nested knowledge schemas. Results Here we present Structured Prompt Interrogation and Recursive Extraction of Semantics (SPIRES), a Knowledge Extraction approach that relies on the ability of Large Language Models (LLMs) to perform zero-shot learning and general-purpose query answering from flexible prompts and return information conforming to a specified schema. Given a detailed, user-defined knowledge schema and an input text, SPIRES recursively performs prompt interrogation against an LLM to obtain a set of responses matching the provided schema. SPIRES uses existing ontologies and vocabularies to provide identifiers for matched elements. We present examples of applying SPIRES in different domains, including extraction of food recipes, multi-species cellular signaling pathways, disease treatments, multi-step drug mechanisms, and chemical to disease relationships. Current SPIRES accuracy is comparable to the mid-range of existing Relation Extraction methods, but greatly surpasses an LLM’s native capability of grounding entities with unique identifiers. SPIRES has the advantage of easy customization, flexibility, and, crucially, the ability to perform new tasks in the absence of any new training data. This method supports a general strategy of leveraging the language interpreting capabilities of LLMs to assemble knowledge bases, assisting manual knowledge curation and acquisition while supporting validation with publicly-available databases and ontologies external to the LLM. Availability and implementation SPIRES is available as part of the open source OntoGPT package: https://github.com/monarch-initiative/ontogpt.

59 BASIC BIOLOGICAL SCIENCES↗

A Simple Standard for Sharing Ontological Mappings (SSSOM)

Abstract Despite progress in the development of standards for describing and exchanging scientific information, the lack of easy-to-use standards for mapping between different representations of the same or similar objects in different databases poses a major impediment to data integration and interoperability. Mappings often lack the metadata needed to be correctly interpreted and applied. For example, are two terms equivalent or merely related? Are they narrow or broad matches? Or are they associated in some other way? Such relationships between the mapped terms are often not documented, which leads to incorrect assumptions and makes them hard to use in scenarios that require a high degree of precision (such as diagnostics or risk prediction). Furthermore, the lack of descriptions of how mappings were done makes it hard to combine and reconcile mappings, particularly curated and automated ones. We have developed the Simple Standard for Sharing Ontological Mappings (SSSOM) which addresses these problems by: (i) Introducing a machine-readable and extensible vocabulary to describe metadata that makes imprecision, inaccuracy and incompleteness in mappings explicit. (ii) Defining an easy-to-use simple table-based format that can be integrated into existing data science pipelines without the need to parse or query ontologies, and that integrates seamlessly with Linked Data principles. (iii) Implementing open and community-driven collaborative workflows that are designed to evolve the standard continuously to address changing requirements and mapping practices. (iv) Providing reference tools and software libraries for working with the standard. In this paper, we present the SSSOM standard, describe several use cases in detail and survey some of the existing work on standardizing the exchange of mappings, with the goal of making mappings Findable, Accessible, Interoperable and Reusable (FAIR). The SSSOM specification can be found at http://w3id.org/sssom/spec. Database URL: http://w3id.org/sssom/spec

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Automated annotation of scientific texts for ML-based keyphrase extraction and validation

Advanced omics technologies and facilities generate a wealth of valuable data daily; however, the data often lack the essential metadata required for researchers to find, curate, and search them effectively. The lack of metadata poses a significant challenge in the utilization of these data sets. Machine learning (ML)–based metadata extraction techniques have emerged as a potentially viable approach to automatically annotating scientific data sets with the metadata necessary for enabling effective search. Text labeling, usually performed manually, plays a crucial role in validating machine-extracted metadata. However, manual labeling is time-consuming and not always feasible; thus, there is a need to develop automated text labeling techniques in order to accelerate the process of scientific innovation. This need is particularly urgent in fields such as environmental genomics and microbiome science, which have historically received less attention in terms of metadata curation and creation of gold-standard text mining data sets. In this paper, we present two novel automated text labeling approaches for the validation of ML-generated metadata for unlabeled texts, with specific applications in environmental genomics. Our techniques show the potential of two new ways to leverage existing information that is only available for select documents within a corpus to validate ML models, which can then be used to describe the remaining documents in the corpus. The first technique exploits relationships between different types of data sources related to the same research study, such as publications and proposals. The second technique takes advantage of domain-specific controlled vocabularies or ontologies. In this paper, we detail applying these approaches in the context of environmental genomics research for ML-generated metadata validation. Our results show that the proposed label assignment approaches can generate both generic and highly specific text labels for the unlabeled texts, with up to 44% of the labels matching with those suggested by a ML keyword extraction algorithm.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

The Coastal Carbon Library and Atlas: Open source soil data and tools supporting blue carbon research and policy

Abstract Quantifying carbon fluxes into and out of coastal soils is critical to meeting greenhouse gas reduction and coastal resiliency goals. Numerous ‘blue carbon’ studies have generated, or benefitted from, synthetic datasets. However, the community those efforts inspired does not have a centralized, standardized database of disaggregated data used to estimate carbon stocks and fluxes. In this paper, we describe a data structure designed to standardize data reporting, maximize reuse, and maintain a chain of credit from synthesis to original source. We introduce version 1.0.0. of the Coastal Carbon Library, a global database of 6723 soil profiles representing blue carbon‐storing systems including marshes, mangroves, tidal freshwater forests, and seagrasses. We also present the Coastal Carbon Atlas, an R‐shiny application that can be used to visualize, query, and download portions of the Coastal Carbon Library. The majority (4815) of entries in the database can be used for carbon stock assessments without the need for interpolating missing soil variables, 533 are available for estimating carbon burial rate, and 326 are useful for fitting dynamic soil formation models. Organic matter density significantly varied by habitat with tidal freshwater forests having the highest density, and seagrasses having the lowest. Future work could involve expansion of the synthesis to include more deep stock assessments, increasing the representation of data outside of the U.S., and increasing the amount of data available for mangroves and seagrasses, especially carbon burial rate data. We present proposed best practices for blue carbon data including an emphasis on disaggregation, data publication, dataset documentation, and use of standardized vocabulary and templates whenever appropriate. To conclude, the Coastal Carbon Library and Atlas serve as a general example of a grassroots F.A.I.R. (Findable, Accessible, Interoperable, and Reusable) data effort demonstrating how data producers can coordinate to develop tools relevant to policy and decision‐making.

Holmquist, James R.↗

The hidden roots of wetland methane emissions

Abstract Wetlands are the largest natural source of methane (CH 4 ) globally. Climate and land use change are expected to alter CH 4 emissions but current and future wetland CH 4 budgets remain uncertain. One important predictor of wetland CH 4 flux, plants, play an important role in providing substrates for CH 4 ‐producing microbes, increasing CH 4 consumption by oxygenating the rhizosphere, and transporting CH 4 from soils to the atmosphere. Yet, there remain various mechanistic knowledge gaps regarding the extent to which plant root systems and their traits influence wetland CH 4 emissions. Here, we present a novel conceptual framework of the relationships between a range of root traits and CH 4 processes in wetlands. Based on a literature review, we propose four main CH 4 ‐relevant categories of root function: gas transport, carbon substrate provision, physicochemical influences and root system architecture. Within these categories, we discuss how individual root traits influence CH 4 production, consumption, and transport (PCT). Our findings reveal knowledge gaps concerning trait functions in physicochemical influences, and the role of mycorrhizae and temporal root dynamics in PCT. We also identify priority research needs such as integrating trait measurements from different root function categories, measuring root‐CH 4 linkages along environmental gradients, and following standardized root ecology protocols and vocabularies. Thus, our conceptual framework identifies relevant belowground plant traits that will help improve wetland CH 4 predictions and reduce uncertainties in current and future wetland CH 4 budgets.

54 ENVIRONMENTAL SCIENCES↗

A starting guide to root ecology: strengthening ecological concepts and standardising root classification, sampling, processing and trait measurements

In the context of a recent massive increase in research on plant root functions and their impact on the environment, root ecologists currently face many important challenges to keep on generating cutting-edge, meaningful and integrated knowledge. Consideration of the below-ground components in plant and ecosystem studies has been consistently called for in recent decades, but methodology is disparate and sometimes inappropriate. This handbook, based on the collective effort of a large team of experts, will improve trait comparisons across studies and integration of information across databases by providing standardised methods and controlled vocabularies. It is meant to be used not only as starting point by students and scientists who desire working on below-ground ecosystems, but also by experts for consolidating and broadening their views on multiple aspects of root ecology. Beyond the classical compilation of measurement protocols, we have synthesised recommendations from the literature to provide key background knowledge useful for: (1) defining below-ground plant entities and giving keys for their meaningful dissection, classification and naming beyond the classical fine-root vs coarse-root approach; (2) considering the specificity of root research to produce sound laboratory and field data; (3) describing typical, but overlooked steps for studying roots (e.g. root handling, cleaning and storage); and (4) gathering metadata necessary for the interpretation of results and their reuse. Most importantly, all root traits have been introduced with some degree of ecological context that will be a foundation for understanding their ecological meaning, their typical use and uncertainties, and some methodological and conceptual perspectives for future research. Considering all of this, we urge readers not to solely extract protocol recommendations for trait measurements from this work, but to take a moment to read and reflect on the extensive information contained in this broader guide to root ecology, including sections I–VII and the many introductions to each section and root trait description. Finally, it is critical to understand that a major aim of this guide is to help break down barriers between the many subdisciplines of root ecology and ecophysiology, broaden researchers’ views on the multiple aspects of root study and create favourable conditions for the inception of comprehensive experiments on the role of roots in plant and ecosystem functioning.

59 BASIC BIOLOGICAL SCIENCES↗

Position-Enhanced Gradient Attack (PEGA) on Medical Language Models

Federated Learning (FL) enables collaborative training of language models on sensitive clinical notes without sharing the data. However, this paradigm is vulnerable to gradient inversion attacks that can reconstruct private data from shared gradients. We find that state-of-the-art attacks are less effective in the medical domain, failing to overcome the unique challenges posed by its specialized vocabulary and unstructured format. To address this, we introduce the Position-Enhanced Gradient Attack (PEGA), a novel attack that makes gradients position-aware by optimizing token and position embeddings simultaneously. PEGA employs two key innovations: a periodic sorting of positional embeddings to resolve token order ambiguity and a late-stage embedding replacement strategy to correct hard-to-recover critical tokens. To evaluate the leakage of sensitive data more directly, we also propose the Unified PHI-Recall (UPHI), a new metric measuring the recovery of Protected Health Information. Experiments on the MIMIC-III dataset show that PEGA significantly outperforms leading attacks like TAG and LAMP, particularly in its ability to reconstruct identifiable patient information, exposing a more severe and nuanced privacy risk in federated medical NLP.

Xu, Nuo [University of Minnesota]↗

NEPATEC2.0: NEPA Text Corpus v2.0

The National Environmental Policy Act of 1969, as amended (NEPA), is a major environmental law in the United States, requiring Federal agencies to consider and document potential environmental impacts before deciding on a proposed action. Modernization of NEPA and permitting processes faces significant challenges due to the lack of standardized formats and interoperable systems for organizing and sharing NEPA-related information across agencies. Much of the information gathered during NEPA reviews is written into documents such as categorical exclusions, environmental assessments, and environmental impact statements, then filed in predominately independent agency file stores that may or may not be publicly accessible. The application of metadata and data standards, such as those recommended by the Council on Environmental Quality (CEQ), to NEPA documents offers a shared vocabulary and structure for key entities like projects, processes, and documents that can streamline information exchange and enhance collaboration across systems. In this work, we publicly release NEPATEC2.0, an expanded corpus of NEPA documents with associated metadata. NEPATEC2.0 encompasses approximately 120,000 documents from 60,000 projects prepared by more than 60 different agencies. Modeled to align with CEQ metadata standards, NEPATEC2.0 promotes consistency in environmental reviews and supports the ongoing effort to modernize permitting technologies by facilitating more transparent, efficient, and data-driven decision-making. Importantly, NEPATEC2.0 demonstrates the possibilities and limitations of large language model-based prompting to extract information from NEPA documents at scale.

environmental review↗

Path-BigBird: An AI-Driven Transformer Approach to Classification of Cancer Pathology Reports

PURPOSE Surgical pathology reports are critical for cancer diagnosis and management. To accurately extract information about tumor characteristics from pathology reports in near real time, we explore the impact of using domain-specific transformer models that understand cancer pathology reports. METHODS We built a pathology transformer model, Path-BigBird, by using 2.7 million pathology reports from six SEER cancer registries. We then compare different variations of Path-BigBird with two less computationally intensive methods: Hierarchical Self-Attention Network (HiSAN) classification model and an offthe-shelf clinical transformer model (Clinical BigBird). We use five pathology information extraction tasks for evaluation: site, subsite, laterality, histology, and behavior. Model performance is evaluated by using macro and micro F 1 scores. RESULTS We found that Path-BigBird and Clinical BigBird outperformed the HiSAN in all tasks. Clinical BigBird performed better on the site and laterality tasks. Versions of the Path-BigBird model performed best on the two most difficult tasks: subsite (micro F 1 score of 72.53, macro F 1 score of 35.76) and histology (micro F 1 score of 80.96, macro F 1 score of 37.94). The largest performance gains over the HiSAN model were for histology, for which a Path-BigBird model increased the micro F 1 score by 1.44 points and the macro F 1 score by 3.55 points. Overall, the results suggest that a Path-BigBird model with a vocabulary derived from wellcurated and deidentified data is the best-performing model. CONCLUSION The Path-BigBird pathology transformer model improves automated information extraction from pathology reports. Although Path-BigBird outperforms Clinical BigBird and HiSAN, these less computationally expensive models still have utility when resources are constrained.

60 APPLIED LIFE SCIENCES↗