Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “publicly-available database”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Structured Prompt Interrogation and Recursive Extraction of Semantics (SPIRES): a method for populating knowledge bases using zero-shot learning

Abstract Motivation Creating knowledge bases and ontologies is a time consuming task that relies on manual curation. AI/NLP approaches can assist expert curators in populating these knowledge bases, but current approaches rely on extensive training data, and are not able to populate arbitrarily complex nested knowledge schemas. Results Here we present Structured Prompt Interrogation and Recursive Extraction of Semantics (SPIRES), a Knowledge Extraction approach that relies on the ability of Large Language Models (LLMs) to perform zero-shot learning and general-purpose query answering from flexible prompts and return information conforming to a specified schema. Given a detailed, user-defined knowledge schema and an input text, SPIRES recursively performs prompt interrogation against an LLM to obtain a set of responses matching the provided schema. SPIRES uses existing ontologies and vocabularies to provide identifiers for matched elements. We present examples of applying SPIRES in different domains, including extraction of food recipes, multi-species cellular signaling pathways, disease treatments, multi-step drug mechanisms, and chemical to disease relationships. Current SPIRES accuracy is comparable to the mid-range of existing Relation Extraction methods, but greatly surpasses an LLM’s native capability of grounding entities with unique identifiers. SPIRES has the advantage of easy customization, flexibility, and, crucially, the ability to perform new tasks in the absence of any new training data. This method supports a general strategy of leveraging the language interpreting capabilities of LLMs to assemble knowledge bases, assisting manual knowledge curation and acquisition while supporting validation with publicly-available databases and ontologies external to the LLM. Availability and implementation SPIRES is available as part of the open source OntoGPT package: https://github.com/monarch-initiative/ontogpt.

59 BASIC BIOLOGICAL SCIENCES↗

Global root traits (GRooT) database

Motivation: Trait data are fundamental to the quantitative description of plant form and function. Although root traits capture key dimensions related to plant responses to changing environmental conditions and effects on ecosystem processes, they have rarely been included in large-scale comparative studies and global models. For instance, root traits remain absent from nearly all studies that define the global spectrum of plant form and function. Thus, to overcome conceptual and methodological roadblocks preventing a widespread integration of root trait data into large-scale analyses we created the Global Root Trait (GRooT) Database. GRooT provides readyto- use data by combining the expertise of root ecologists with data mobilization and curation. Specifically, we (a) determined a set of core root traits relevant to the description of plant form and function based on an assessment by experts, (b) maximized species coverage through data standardization within and among traits, and (c) implemented data quality checks. Main types of variables contained: GRooT contains 114,222 trait records on 38 continuous root traits. Spatial location and grain: Global coverage with data from arid, continental, polar, temperate and tropical biomes. Data on root traits were derived from experimental studies and field studies. Time period and grain: Data were recorded between 1911 and 2019. Major taxa and level of measurement: GRooT includes root trait data for which taxonomic information is available. Trait records vary in their taxonomic resolution, with subspecies or varieties being the highest and genera the lowest taxonomic resolution available. It contains information for 184 subspecies or varieties, 6,214 species, 1,967 genera and 254 families. Owing to variation in data sources, trait records in the database include both individual observations and mean values. Software format: GRooT includes two csv files. A GitHub repository contains the csv files and a script in R to query the database.

59 BASIC BIOLOGICAL SCIENCES↗

Building Performance Database API (BPD API) v2.1

The Building Performance Database (BPD) is the largest publicly-available source of measured energy performance data for buildings in the United States. It contains information about the building's energy use, location, and physical and operational characteristics. The BPD can be used by building owners, operators, architects and engineers to compare a building's energy efficiency against customized peer groups, identify energy efficiency opportunities, and set energy efficiency targets. It can also be used by energy efficiency program implementers and policymakers to analyze energy efficiency features and trends in the building stock. The BPD compiles data from various data sources, converts it into a standard format, cleanses and quality checks the data, and provides users with access to the data in a way that maintains anonymity for data providers. This software is the database and the Application Programming Interface (API). Users can utilize the BPD's data to develop their own applications using the API. Version 2.1 included a major update for multiple years of data and refactoring of code for faster queries.

Mathew, Paul↗

A Novel Framework for Performance Evaluation and Design Optimization of PCM Embedded Heat Exchangers for the Built Environment

This research sheds light on the performance evaluation and design optimization of PCM-HXs for the built environment, addressing several barriers to practical issues to PCM-HX commercialization such as modeling aspects (i.e., modeling expertise and computational / time investment, etc.), manufacturing aspects (i.e., at-scale manufacturing, cost assessments, etc.) and experimental performance assessment (i.e., reliable experimental data, assessment of multiple PCM-working fluid combinations, etc.). We present a novel, comprehensive, and experimentally-validated design optimization framework for PCM-HXs capable of simulating any PCM-HX geometry with reasonable accuracy and significant computational time savings when compared to traditional CFD-based design practices. The framework was validated for a wide range of PCM-HX configurations, including a design optimization for a domestic hot water heater application where TES partially replaces electrical heating input. The resulting PCM-HXs were found to deliver 34-68% of the total daily hot water supply with only 5-10% package volume increase from the water heater, thus within U.S. DOE targets for TES systems. To identify the most promising HXs for PCM applications, first-order geometry and cost analyses were conducted based on off-the-shelf HX products. As part of this work, 9 PCM-HX prototypes were manufactured using additive and conventional manufacturing methods. Detailed economy-of-scale assessments were conducted for the most promising PCM-HXs and were found to have a good outlook for the next 5-10 years. The PCM-HX design optimization framework was validated through comprehensive in-house experimental testing using newly-developed PCM-to-fluid test facilities. In total,10 total in-house component-level experiments were conducted using these prototypes, including 9 with water and 1 with refrigerant (R410A) as the working fluid. It was found that the framework can successfully predict experimental thermal-hydraulic performance within ±10-20% the first time without manual design changes, eliminating the need for time-consuming and expensive prototyping efforts as part of the design process. As part of this work, a publicly-available PCM web tool was released which includes a PCM property database (531 PCMs) and PCM-HX modeling tool to assist the design community on common PCM-HX use-cases, e.g., single/multiple flow path(s) fluid-to-PCM and air-to-fluid-to-PCM configurations (https://ceeeweb.umd.edu/pcmapp/). This work will accelerate the design and time to market for next generation PCM-HXs.

25 ENERGY STORAGE↗

High-Temperature Gas-Cooled Reactor Research Survey and Overview: Preliminary Data Platform Construction for the Nuclear Energy University Program

Since the U.S. Department of Energy Office of Nuclear Energy initiated the Nuclear Energy University Program (NEUP) in 2009, there are 29 NEUP projects focusing on high-temperature gas-cooled reactor (HTGR) research up to July 2022. The resultant research product, either experimental or computational, were published as final NEUP reports, journal articles and conference proceedings. However, these federally funded products have been scattered and sometimes cannot be easily accessed. To improve access to this valuable HTGR validation data and optimize the return on the significant investment made by the Department of Energy, the Advanced Reactor Technologies (ART) Gas-Cooled Reactor (GCR) program started a survey of completed and ongoing HTGR NEUP projects to develop a public-access database specific for HTGRs applications that can be used to retrieve computational fluid dynamics and system code validation data. This effort will help guide future NEUP-funded research, define new state of the ART Phenomena Identification and Ranking Table (PIRT), and promote the usage of this data in the codes validation matrices. This report provides an overview of the NEUP-funded HTGR-related research projects from Fiscal Year (FY) 2009–2021 and identifies validation knowledge gaps still existing in HTGR thermal-fluid research. A preliminary data platform has been developed for the 29 NEUP projects investigating HTGR thermal hydraulics, including their final reports as well as the available scientific publications. As an ultimate goal for this work, the ART-GCR program will create a central database at Idaho National Laboratory to identify, organize, and store these datasets generated by experimental investigations or computational models, experimental facility descriptions, and publicly-available academic products from the HTGR-related NEUP projects and provide future guidance for the storage and transmission of important project documentations for later NEUP projects as well.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Pennsylvania Department of Environmental Protection (PA DEP) 26r Detailed Produced Water Compositions (version 1.0)

A database of geochemical compositions of aqueous species in produced water reported to the PA DEP. Samples were collected between mid-2012 to early-2020. Data from publicly-available PA DEP 26r reports were scraped from pdf files and cumulated into tabular spreadsheet format for >1000 produced water streams from Marcellus wells in Pennsylvania. In addition to providing the original values, the NETL NEWTS team has reformatted the dataset to allow sample streams to be easily copied into OLI Studio and Geochemist WorkBench (GWB) software for modeling the geochemistry and the recovery of critical minerals, such as lithium, from these produced water streams. In addition, a version of the dataset has been included with predictions for some missing values in the original dataset using machine learning techniques within CoDaRT software, a public ML software developed by the Nation Energy Technology Laboratory. We have made the Input into CoDaRT and one example output from CoDaRT available in this dataset.

Aqueous Chemistry↗

Optimal CO2 Transport and Storage Cost Screening: Application Example

Poster on “Optimal CO2 Transport and Storage Cost Screening: Application Example” for the CCUS 2025 conference held in Houston, Texas March 3-5, 2025. A major challenge to commercial scale CCS deployment from the perspective of coal and natural gas-fired power plants is understanding cost-optimal CO2 transport and viable geologic storage options. This study demonstrates unique workflows, using NETL-developed, publicly-available models and tools, to efficiently estimate optimal CO2 transport and storage (T&S) costs for each of the CO2 sources in NETL’s Carbon Capture Retrofit Databases (CCRD) for Electricity Generating Units. The results demonstrate the impact of cost-drivers on optimal T&S, and trends in optimal T&S data, based on real point sources that could be retrofitted with CO2 source technologies.

application example↗