Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “keyword”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Understanding Event Trajectories Across Massive Temporal Datasets with Word Embeddings and Visualization

In collaboration with researchers from Virginia Tech, Savannah River National Laboratory has continued development of a natural language processing pipeline to identify and extract events of interest from massive open data sources in the domain of worldwide state-sponsored civil nuclear energy. The foundation of the pipeline is built on compass aligned temporal word embedding models, whereby contextual shifts are automatically identified by comparing keyword embedding vectors across successive time windows. Within the approach, a contextual shift indicates the occurrence of a potential event of interest. However, in such a broad topical domain that captures events at a global scale, across various life cycle stages, and across numerous different technology types, a user that is monitoring events may have broad interests in capturing many different event types with varying degrees of signal. As such, the quantity of information that may be returned from an automated event extraction pipeline can be substantial, requiring manual effort to sift through the information to identify any relevant bits of information. Therefore, a more streamlined workflow that aids in directing a user toward specific information at different points in time is necessary. The workflow presented here has been developed with this concept in mind, built on top of the initial prototype event extraction pipeline, whereby a user can analyze temporal text-based data sources at multiple different contextual levels to isolate key points in time and key subdomains captured within a data corpus. Using multiple corpuses that consist of approximately 7 million Tweets and 7 million news articles, the team has extended compass aligned temporal word embedding models to establish an interconnected and hierarchical structure that relates known key words of interest to documents, local topics (i.e., within a time window), and global topics across the corpuses. All of this information is packaged into a visual analytics system that is linked to the information extraction pipeline and enables a user to identify contextual information that describes the evolution of a high dimensional embedding space across time to isolate changes of interest and explore associated events. This report demonstrates the use of these analytics and a means to fuse information across multiple datasets.

97 MATHEMATICS AND COMPUTING↗

ITS Version 6.7: The Integrated TIGER Series of Coupled Electron/Photon Monte Carlo Transport Codes User's Manual

ITS is a powerful software package permitting state-of-the-art Monte Carlo solution of linear time-independent coupled electron/photon radiation transport problems, with or without the presence of macroscopic electric and magnetic fields of arbitrary spatial dependence. Our goal has been to simultaneously maximize operational simplicity and physical accuracy. Through a set of preprocessor directives, the user selects one of the many ITS codes. The ease with which the make system is applied combines with an input scheme based on order-independent descriptive keywords that makes maximum use of defaults and internal error checking to provide experimentalists and theorists alike with a method for the routine but rigorous solution of sophisticated radiation transport problems. Physical rigor is provided by employing accurate cross sections, sampling distributions, and physical models for describing the production and transport of the electron/photon cascade from 1.0 GeV down to 1.0 keV. The availability of source code permits the more sophisticated user to tailor the codes to specific applications and to extend the capabilities of the codes to more complex applications. Version 6, the latest version of ITS, contains (1) improvements to the ITS 5.0 codes, and (2) conversion to Fortran 95. The general user friendliness of the software has been enhanced through memory allocation to reduce the need for users to modify and recompile the code.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Utilizing Ontology Structures To Curate the DOE-NETL Carbon Storage Open Database

The specialized ontology for the Carbon Storage Open Database will enable more rapid assignment of appropriate symbology standards for visualization improvements, optimize topical and spatial tagging within keywords, and improve flexibility for utilization in existing data repositories such as EDX. This effort also aims to establish a foundation for utilization of ontologies for organization of other data related to geologic carbon storage in the future.

Martin, Abigail↗

Assessing the Current State of U.S. Energy Equity Regulation and Legislation (2024 Update) [Slides]

Lawrence Berkeley National Laboratory (LBNL) partnered with E9 Insight (E9) to update the 2022 database of executive, legislative, and regulatory actions focusing on energy equity and directed at electricity and natural gas utilities. The resulting database contains 307 energy equity actions, which consist of documents (e.g., bills, dockets, and executive orders) identified through keyword searches associated with energy equity. Based on our review, almost half of states (32 + DC) were taking some sort of action on energy equity (i.e., executive order, PUC activity, agency plan, or executive bill). In the report, we explored temporal trends to track state actions over time across regions, outcomes, equity tenets, and objectives. We also explored what drivers led to what outcomes. Drivers were organized into legislative, regulatory, executive, and stakeholder-driven.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Datasets for Custom-trained Machine-learning Interatomic Potentials: Nitric Acid Aqueous Solution

This dataset was generated using an iterative active learning strategy with the ArcaNN software package (https://github.com/arcann-chem/arcann_training) to train machine-learning interatomic potentials (MLIPs) for aqueous nitric acid. Each active-learning cycle consisted of three stages: (1) training, (2) exploration, and (3) labeling. The initial training set comprised approximately 800 randomly selected configurations from a previous study by Lewis et al. (https://doi.org/10.1021/jp205510q), which investigated nitric acid solutions at 2, 3, 4, and 5 mol/L. For all configurations, single-point calculations of atomic forces and total energies were performed at the quantum density functional theory BLYP-D2 and PBE-D3 levels of theory using the CP2K Quickstep module. Valence electrons were treated explicitly, while core electrons on all atoms were represented by norm-conserving Goedecker–Teter–Hutter (GTH) pseudopotentials. Long-range dispersion interactions were accounted for using Grimme dispersion corrections. Wave functions were expanded in a mixed Gaussian-and-plane-wave scheme using TZV2P-MOLOPT basis sets for all elements and an 800 Ry auxiliary plane-wave cutoff for the electron density. Self-consistent field convergence was accelerated using orbital transformation and Direct Inversion in the Iterative Subspace, with a convergence threshold of 10^{-6}. All single-point calculations were carried out in periodic orthorhombic cells whose dimensions match those of the molecular configurations sampled from earlier trajectories. The CELL_REF keyword in CP2K was used to define a fixed reference cell, ensuring consistency in the reference data used for MLIP training, particularly when cell fluctuations are present in NpT simulations. The resulting high-fidelity energies and forces constitute the ground-truth labels used to train the MLIPs contained in this dataset.

Dinpajooh, Mohammadhasan [Pacific Northwest Nation↗

Custom-trained Machine-learning Interatomic Potentials: ZnCl2 Aqueous Solution

This dataset was generated using an iterative active-learning strategy implemented in the ArcaNN software package (https://github.com/arcann-chem/arcann_training) to train machine-learning interatomic potentials for aqueous ZnCl2 solutions. Each active-learning cycle consisted of three stages: training, exploration, and labeling. The initial training set combined configurations generated in this work from enhanced-sampling ab initio molecular dynamics simulations with configurations from a previously reported neural-network-potential study of aqueous ZnCl2. The enhanced-sampling ab initio molecular dynamics simulations involved Zn–Cl separation and the chloride coordination number around Zn²? as collective variables. These configurations served as the seed dataset. Subsequent active-learning cycles expanded the training set by identifying and labeling configurations that were poorly represented by the current models, thereby improving coverage of ion-association states and changes in local coordination and charge-state environments relevant to the solution free-energy landscape. For all selected configurations, single-point calculations of the total energies and atomic forces were performed within density functional theory using the CP2K Quickstep module. Reference calculations employed the revPBE-D3 and r2SCAN exchange-correlation functionals. Motivated by recent work on aqueous Zn²?, the main revPBE calculations omitted D3 dispersion contributions involving Zn²?, while retaining the D3 correction for water and chloride. For comparison, fully dispersion-corrected revPBE-D3 reference calculations were also performed, with D3 applied to all species, including Zn²?. Valence electrons were treated explicitly, while core electrons were represented using norm-conserving Goedecker–Teter–Hutter pseudopotentials. The wave functions were expanded using the mixed Gaussian-and-plane-wave scheme with TZV2P-MOLOPT basis sets for all elements and a 600 Ry auxiliary plane-wave cutoff for the electron density. Self-consistent-field convergence was accelerated using the orbital-transformation and Direct Inversion in the Iterative Subspace algorithms, with a convergence threshold of 10?6. All single-point calculations were performed in periodic orthorhombic cells. The CELL_REF keyword in CP2K was used to define a fixed reference cell with a box length of 25 Å. This treatment ensured a consistent reference for configurations extracted from NpT trajectories with fluctuating cell dimensions. The resulting DFT energies and atomic forces constitute the ground-truth labels used to train the MLIPs. The resulting MLIP was trained for aqueous ZnCl2 solutions spanning concentrations from 0 to 30 molal and a broad pH range, from strongly acidic to strongly basic conditions. Representative examples of configurations included in the MLIP training dataset are provided below. These include 1) Representative configurations from the dataset labeled at the revPBE-D3 level, with D3 dispersion interactions involving Zn2+ excluded (revPBE-wo-D3). 2) Representative configurations from the dataset labeled at the fully dispersion-corrected revPBE-D3 level, with D3 interactions applied to all species, including Zn2+ (revPBE-D3). 3) Representative configurations from the dataset labeled at the r2SCAN level of theory (r2SCAN).

Dinpajooh, Mohammadhasan [Pacific Northwest Nation↗

Temporal partitioning of hatching, maturation, and surface activity by reptiles in Florida longleaf pine-wiregrass sandhills

Temporal partitioning of life history traits among syntopic reptiles can facilitate co-occurrence, but may be influenced by environmental factors and evolutionary history. We used 24 years of continuous capture data in the Florida sandhills to evaluate the timing and duration of hatching, maturation and/or surface activity for ten reptile species, spanning multiple clutch strategies, taxonomic relationships, and habits. We hypothesised: i) species would differ in seasonal timing of hatching and maturation; ii) hatching and maturation periods would be more seasonally-synchronised in fossorial than terrestrial or semi-aquatic reptiles; iii) monthly and annual temperature anomalies would be positively related to hatching, maturation, and surface activity anomalies, and iv) groupings of reptiles by clutch strategy, taxonomic relationship, and habit, would explain more variation in the timing and duration of hatching and maturation than species alone. Seasonal timing of response variables varied widely among species. Hatching peaked for > 1 species during most calendar months. Maturation and surface activity periods ranged from aseasonal to highly-seasonal among species. Hatching began 1.5 months earlier and was more prolonged for terrestrial than fossorial species overall. Hatching peaked in early to mid-summer for terrestrial and fossorial species, and winter for the semi-aquatic Kinosternon subrubrum. Terrestrial and fossorial species did not differ in average timing, duration, or overlap of maturation periods; semi-aquatic Liodytes pygaea matured more consistently across all seasons than other species. Monthly temperature anomalies were negatively correlated with monthly maturation for Plestiodon egregius. Annual temperature and precipitation anomalies were related to annual hatching, maturation, and surface activity trends for several species. Taxonomic relationship, habit, and species explained some variation in hatching and maturation timing and duration. Our results illustrate the influence of environment and evolutionary relationships on the timing of important life history traits. Keywords: Age class, Drift fence, Community, Interspecific, Life history

Zoology↗

MaizeMine: A Data Mining Warehouse for the Maize Genetics and Genomics Database

MaizeMine is the data mining resource of the Maize Genetics and Genome Database (MaizeGDB; http://maizemine.maizegdb.org). It enables researchers to create and export customized annotation datasets that can be merged with their own research data for use in downstream analyses. MaizeMine uses the InterMine data warehousing system to integrate genomic sequences and gene annotations from the Zea mays B73 RefGen_v3 and B73 RefGen_v4 genome assemblies, Gene Ontology annotations, single nucleotide polymorphisms, protein annotations, homologs, pathways, and precomputed gene expression levels based on RNA-seq data from the Z. mays B73 Gene Expression Atlas. MaizeMine also provides database cross references between genes of alternative gene sets from Gramene and NCBI RefSeq. MaizeMine includes several search tools, including a keyword search, built-in template queries with intuitive search menus, and a QueryBuilder tool for creating custom queries. The Genomic Regions search tool executes queries based on lists of genome coordinates, and supports both the B73 RefGen_v3 and B73 RefGen_v4 assemblies. The List tool allows you to upload identifiers to create custom lists, perform set operations such as unions and intersections, and execute template queries with lists. When used with gene identifiers, the List tool automatically provides gene set enrichment for Gene Ontology (GO) and pathways, with a choice of statistical parameters and background gene sets. With the ability to save query outputs as lists that can be input to new queries, MaizeMine provides limitless possibilities for data integration and meta-analysis.

59 BASIC BIOLOGICAL SCIENCES↗

Image-Guided Intraoperative Assessment of Surgical Margins in Oral Cavity Squamous Cell Cancer: A Diagnostic Test Accuracy Review

The assessment of resection margins during surgery of oral cavity squamous cell cancer (OCSCC) dramatically impacts the prognosis of the patient as well as the need for adjuvant treatment in the future. Currently there is an unmet need to improve OCSCC surgical margins which appear to be involved in around 45% cases. Intraoperative imaging techniques, magnetic resonance imaging (MRI) and intraoral ultrasound (ioUS), have emerged as promising tools in guiding surgical resection, although the number of studies available on this subject is still low. The aim of this diagnostic test accuracy (DTA) review is to investigate the accuracy of intraoperative imaging in the assessment of OCSCC margins. By using the Cochrane-supported platform Review Manager version 5.4, a systematic search was performed on the online databases MEDLINE-EMBASE-CENTRAL using the keywords “oral cavity cancer, squamous cell carcinoma, tongue cancer, surgical margins, magnetic resonance imaging, intraoperative, intra-oral ultrasound”. Ten papers were identified for full-text analysis. The negative predictive value (cutoff < 5 mm) for ioUS ranged from 0.55 to 0.91, that of MRI ranged from 0.5 to 0.91; accuracy analysis performed on four selected studies showed a sensitivity ranging from 0.07 to 0.75 and specificity ranging from 0.81 to 1. Image guidance allowed for a mean improvement in free margin resection of 35%. IoUS shows comparable accuracy to that of ex vivo MRI for the assessment of close and involved surgical margins, and should be preferred as the more affordable and reproducible technique. Both techniques showed higher diagnostic yield if applied to early OCSCC (T1–T2 stages), and when histology is favorable.

60 APPLIED LIFE SCIENCES↗

Beneficial Use Impairments, Degradation of Aesthetics, and Human Health: A Review

In environmental programs and blue/green space development, improving aesthetics is a common goal. There is broad interest in understanding the relationship between ecologically sound environments that people find aesthetically pleasing and human health. However, to date, few studies have adequately assessed this relationship, and no summaries or reviews of this line of research exist. Therefore, we undertook a systematic literature review to determine the state of science and identify critical needs to advance the field. Keywords identified from both aesthetics and loss of habitat literature were searched in PubMed and Web of Science databases. After full text screening, 19 studies were included in the review. Most of these studies examined some measure of greenspace/bluespace, primarily proximity. Only one study investigated the impacts of making space quality changes on a health metric. The studies identified for this review continue to support links between green space and various metrics of health, with additional evidence for blue space benefits on health. No studies to date adequately address questions surrounding the beneficial use impairment degradation of aesthetics and how improving either environmental quality (remediation) or ecological health (restoration) efforts have impacted the health of those communities.

60 APPLIED LIFE SCIENCES↗

Procedure Parsing: A Method for Parsing Handwritten Documents into Computer-Based Procedures

The nuclear industry is heavily procedure driven, where almost everything has a step-by-step instruction that is expected to be followed in detail. Historically, these procedures were printed on paper copies. Recently, the industry transitioned towards electronic copies (i.e., PDFs on tablets). One major drive for this transition is the introduction of human error and loss of situation awareness when using paper copies. However, electronic copies of documents inherently have the same error traps as their paper cousins. Therefore, there is an increased interest in a way to utilize the information in the step-by-step guidance, but to present it in a dynamic manner that guides the user and adapts to any encountered conditions. Researchers at Idaho National Laboratory propose a flexible, automated method based on document parsing and augmented by natural language processing (NLP) techniques, to address these shortcomings and capitalize on these recent advancements in machine learning. The proposed method provides a cost-effective solution for computer-assisted procedure parsing of hand-written control room procedures, originally authored in Word or PDF formats, into instructions that can be displayed as computer-based procedures (CBP) in a modern graphical user interface. The researchers devised, implemented and demonstrated the Operating Procedure Extender for Novel Systems (OPENS) method in 2020. The key to OPENS is to map the original procedure text into a context-free grammar, tying content to equipment, locations, and other steps, actions, etc. This formal grammar is then used to isolate and define keywords and actions verbs, such as “measure” or “evaluate” and tie them to specific equipment referenced within that step or located in other steps, substeps, actions, subactions and tables throughout the procedure. OPENS generates an abstract syntax tree from the document which it uses to store a copy of this information in the open-standard, machine-readable and human-readable file formats XML and JSON. The XML is useful to preserve the relational aspects of the procedure for referencing tables and branching information so the user can be directed to the next appropriate active step based on the values entered for that step and previous steps. The JSON is useful for storing and exchanging data objects used to track responses to previous steps and state changes in simulated environments. In future iterations, these formats can also be used for storing more detailed information about input during plant operation or simulation. The techniques the researcher developed could further be improved by integration of recent advancements in machine learning. NLP methods could standardize documents, correct for grammatical error, and provide automated semantic validation. The researcher expects that self-supervised techniques applied to collections of natural language instructions could strengthen the model with broader context. All these methods together give us a practical way to automatically extract protocols from documents and user interactions, empowering researchers, procedure writers and nuclear operators while moving the industry forward.

99 GENERAL AND MISCELLANEOUS↗

JSONize: A Scalable Machine Learning Pipeline to Model Medical Notes as Semi-structured Documents

The Department of Veteran's Affairs (VA) archives the largest corpora of clinical notes in their corporate data warehouse (CDW) as unstructured text data. Unstructured text easily supports keyword searches and regular expressions. Often these simple searches do not adequately support the complex searches that need to be performed on notes. For example, a researcher may want all notes with a Duke Treadmill Score less than 5 or people that smoke more than 1 pack per day. Range queries like this and more can be supported by modelling text as semi-structured documents. In this paper, we implement a scalable machine learning pipeline that models plain medical text as useful semi-structured documents. We improve on existing models and achieve a F1-score of 0.912 and scale our methods to the entire VA corpus.

Rush III, Everett↗

Can machine learning predict fuel properties accurately?

High-potential molecules derived from biomass sources may suitably replace or supplement traditional nonrenewable hydrocarbon fuels to reduce pollution and fuel processing cost. Experimental property testing of these bioproducts is usually conducted years after initial bench-scale experiments, due to high experimental costs and/or high volume requirements. However, neglecting to conduct property testing early in the pathway development cycle can lead to investments spent on scaling-up production of bioproducts and biofuels that do not perform as expected. Instead, machine-learning techniques can be used to develop quantitative structure–property relationships for molecules using a relatively large training set of molecular descriptor data. For this study, we compiled measured properties, IR spectra, and molecular descriptors of bio-based molecules from databases and published studies for training models of bioproduct properties. We trained regression models with molecular descriptors and will compare results of different estimators. This study describes the first steps towards a performance prediction tool for bio-based alternative fuels. Keywords: Machine learning, biofuels, jet fuels, fuel properties

Mayer, Morgan A.↗

Open-source FPGA-ML codesign for the MLPerf Tiny Benchmark

We present our development experience and recent results for the MLPerf Tiny Inference Benchmark on field-programmable gate array (FPGA) platforms. We use the open-source hls4ml and FINN workflows, which aim to democratize AI-hardware codesign of optimized neural networks on FPGAs. We present the design and implementation process for the keyword spotting, anomaly detection, and image classification benchmark tasks. The resulting hardware implementations are quantized, configurable, spatial dataflow architectures tailored for speed and efficiency and introduce new generic optimizations and common workflows developed as a part of this work. The full workflow is presented from quantization-aware training to FPGA implementation. The solutions are deployed on system-on-chip (Pynq-Z2) and pure FPGA (Arty A7-100T) platforms. The resulting submissions achieve latencies as low as 20 $\mu$s and energy consumption as low as 30 $\mu$J per inference. We demonstrate how emerging ML benchmarks on heterogeneous hardware platforms can catalyze collaboration and the development of new techniques and more accessible tools.

Borras, Hendrik↗

VizBrick: A GUI-based Interactive Tool for Authoring Semantic Metadata for Building Datasets

Brick ontology is a unified semantic metadata schema to address the stand-ardization problem of buildings' physical, logical, and virtual assets and the relationships between them. Creating a Brick model for a building dataset means that the dataset's contents are semantically described using the standard terms defined in the Brick ontology. It will enable the benefits of data standardization, without having to recollect or reorganize the data and opens the possibility of automation leveraging the machine readability of the semantic metadata. The problem is that authoring Brick models for building datasets often requires knowledge of semantic technology (e.g., on-tology declarations and RDF syntax) and leads to repeated manual trial and error processes, which can be time-consuming and challenging to do with-out an interactive visual representation of the data. We developed VizBrick, a tool with a graphical user interface that can assist users in creating Brick models visually and interactively without having to understand the Re-source Description Framework (RDF) syntax. VizBrick provides handy ca-pabilities such as keyword search for easy find of relevant brick concepts and relations to their data columns and automatic suggestions of concept mapping. In this demonstration, we present a use-case of VizBrick to show-case how a Brick model can be created for a real-world building dataset.

Lee, Sangkeun (Matt)↗

Radiological Recovery Logistics Tool - 20161

Argonne is building and testing a tool, the Radiological Recovery Logistics Tool (RRLT), that can be used during the response and recovery from a radiological or nuclear incident to effectively allocate appropriate commercial and public works equipment to mitigate, remove, and contain radiological contamination. The requirements for this tool - as well as development of the resulting software - is overseen by a steering committee of stakeholders from DHS's National Urban Security Technology Laboratory (NUSTL), the Federal Emergency Management Agency (FEMA), and the Environmental Protection Agency (EPA). One essential requirement is for RRLT to support the efficient and appropriate allocation of resources for a radiological response. Subsequent discussions between ANL and stakeholders have solidified the nature of this support to include identification of the types of resources to be allocated. The study reported in this paper has both factored fundamental concepts and connections out of this identification process and created a Knowledge Base detailing support goals, response scenarios, and efficacy information on dozens of equipment types. In short, RRLT will dynamically apply these findings to situational conditions surrounding contamination incidents. RRLT's Domain, the model of elements, ideas and relationships with which the tool will work, draws concepts from technical reports and stakeholder vocabularies to connect response goals and scenarios to types of equipment that offer utility towards those goals in those scenarios. RRLT's Knowledge Base will contain details on dozens of equipment types and facilitate the operator's discovery and consumption of these details most pertinent to a dynamically selected subset of goals. The core of its Domain Model is based on a report authored by this team. This report [1] contains a comprehensive list of proposed equipment to accomplish various missions or scenarios that might arise after a large-scale radiological contamination incident in an urban environment or critical infrastructure. The report divides potential response and recovery efforts into five support goals: Survey and monitoring of the contaminated area; Mitigation of received dose to first responders: Decontamination (gross and final) of buildings, vehicles, roadways, parks, and other surfaces: Waste management of solid waste generated during recovery operations: and Containment of wastewater and other waste generated during the response and recovery phases. RRLT's development is driven by use cases. A use case is an intention with which a user approaches the software. Use cases are grouped into delivery increments to schedule development, testing, and presentation to stakeholders. This model partitions the system into seven increments: User Arrival and Authentication, Search and Navigation, Equipment Recommendation, Plan Management, Content Management, and Expanded Access. Once a user 15 authenticated, RRLT will present the user with a dashboard that allows them to explore or search RRLT's content. The dashboard will also include a 'Plan' panel for collecting decisions and relevant observations about an incident at hand to facilitate development of an equipment list. RRLT will offer three general modes of access to items in the knowledge base: - Keyword search for direct discovery of items, - Navigation along predetermined paths from recovery goal towards equipment types, and - Interactive guidance towards equipment types by an autonomous software agent: the Equipment Recommendation Wizard. This presentation will detail progress in the development of the RRLT and also discuss opportunities for those interested in providing feedback on its content and functionality. (authors)

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Empowering Geothermal Research: The Geothermal Data Repository's New AI Research Assistant: Preprint

The Department of Energy's (DOE) Geothermal Data Repository (GDR) team has integrated a Large Language Model (LLM) with the metadata and supporting documents associated with GDR datasets to create an Artificially Intelligent (AI) research assistant. By leveraging work done to make GDR metadata machine-readable and an open-source LLM integration model called the Energy Language Model, developed by the National Renewable Energy Laboratory, AskGDR serves as a virtual research assistant to GDR users. It provides answers to a variety of user-provided questions using natural language processing and generative machine learning. Users can get answers to questions about specific datasets, including inquiries about the equipment, assumptions and methodologies used in the origination of the data; or more abstract questions, such as the applicability of data to specific research fields. AskGDR improves the discoverability of geothermal data by helping guide users to datasets beyond simple keyword searches. It enables users to find data based on properties of the data, discover information contained within supporting documents, and explore data from projects related to their research objectives.

access↗

Empowering Geothermal Research: The Geothermal Data Repository's New AI Research Assistant

The Department of Energy's (DOE) Geothermal Data Repository (GDR) team has integrated a Large Language Model (LLM) with the metadata and supporting documents associated with GDR datasets to create an Artificially Intelligent (AI) research assistant. By leveraging work done to make GDR metadata machine-readable and an open-source LLM integration model called the Energy Language Model, developed by the National Renewable Energy Laboratory, AskGDR serves as a virtual research assistant to GDR users. It provides answers to a variety of user-provided questions using natural language processing and generative machine learning. Users can get answers to questions about specific datasets, including inquiries about the equipment, assumptions and methodologies used in the origination of the data; or more abstract questions, such as the applicability of data to specific research fields. AskGDR improves the discoverability of geothermal data by helping guide users to datasets beyond simple keyword searches. It enables users to find data based on properties of the data, discover information contained within supporting documents, and explore data from projects related to their research objectives. This paper will outline the development, integration, output, and efficacy of the AskGDR LLM, including adherence to scientific rigor through improvements designed to increase the accuracy of generated answers, avoid speculation, and provide proper references for all resources used.

access↗