Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data repository”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Simple Scattering: Lipid nanoparticle structural data repository

Lipid nanoparticles (LNPs) are being intensively researched and developed to leverage their ability to safely and effectively deliver therapeutics. To achieve optimal therapeutic delivery, a comprehensive understanding of the relationship between formulation, structure, and efficacy is critical. However, the vast chemical space involved in the production of LNPs and the resulting structural complexity make the structure to function relationship challenging to assess and predict. New components and formulation procedures, which provide new opportunities for the use of LNPs, would be best identified and optimized using high-throughput characterization methods. Recently, a high-throughput workflow, consisting of automated mixing, small-angle X-ray scattering (SAXS), and cellular assays, demonstrated a link between formulation, internal structure, and efficacy for a library of LNPs. As SAXS data can be rapidly collected, the stage is set for the collection of thousands of SAXS profiles from a myriad of LNP formulations. In addition, correlated LNP small-angle neutron scattering (SANS) datasets, where components are systematically deuterated for additional contrast inside, provide complementary structural information. The centralization of SAXS and SANS datasets from LNPs, with appropriate, standardized metadata describing formulation parameters, into a data repository will provide valuable guidance for the formulation of LNPs with desired properties. To this end, we introduce Simple Scattering, an easy-to-use, open data repository for storing and sharing groups of correlated scattering profiles obtained from LNP screening experiments. Here, we discuss the current state of the repository, including limitations and upcoming changes, and our vision towards future usage in developing our collective knowledge base of LNPs.

59 BASIC BIOLOGICAL SCIENCES↗

Empowering Geothermal Research: The Geothermal Data Repository's New AI Research Assistant

The Department of Energy's (DOE) Geothermal Data Repository (GDR) team has integrated a Large Language Model (LLM) with the metadata and supporting documents associated with GDR datasets to create an Artificially Intelligent (AI) research assistant. By leveraging work done to make GDR metadata machine-readable and an open-source LLM integration model called the Energy Language Model, developed by the National Renewable Energy Laboratory, AskGDR serves as a virtual research assistant to GDR users. It provides answers to a variety of user-provided questions using natural language processing and generative machine learning. Users can get answers to questions about specific datasets, including inquiries about the equipment, assumptions and methodologies used in the origination of the data; or more abstract questions, such as the applicability of data to specific research fields. AskGDR improves the discoverability of geothermal data by helping guide users to datasets beyond simple keyword searches. It enables users to find data based on properties of the data, discover information contained within supporting documents, and explore data from projects related to their research objectives. This paper will outline the development, integration, output, and efficacy of the AskGDR LLM, including adherence to scientific rigor through improvements designed to increase the accuracy of generated answers, avoid speculation, and provide proper references for all resources used.

access↗

Empowering Geothermal Research: The Geothermal Data Repository's New AI Research Assistant: Preprint

The Department of Energy's (DOE) Geothermal Data Repository (GDR) team has integrated a Large Language Model (LLM) with the metadata and supporting documents associated with GDR datasets to create an Artificially Intelligent (AI) research assistant. By leveraging work done to make GDR metadata machine-readable and an open-source LLM integration model called the Energy Language Model, developed by the National Renewable Energy Laboratory, AskGDR serves as a virtual research assistant to GDR users. It provides answers to a variety of user-provided questions using natural language processing and generative machine learning. Users can get answers to questions about specific datasets, including inquiries about the equipment, assumptions and methodologies used in the origination of the data; or more abstract questions, such as the applicability of data to specific research fields. AskGDR improves the discoverability of geothermal data by helping guide users to datasets beyond simple keyword searches. It enables users to find data based on properties of the data, discover information contained within supporting documents, and explore data from projects related to their research objectives.

access↗

Lessons Learned from AskGDR: Usage and Impact Analysis of the Geothermal Data Repository's AI Research Assistant: Preprint

In October of 2024, the Department of Energy's (DOE) Geothermal Data Repository (GDR) team officially launched AskGDR, an AI research assistant resulting from the integration of a Large Language Model (LLM) with the metadata and supporting documents associated with GDR datasets. AskGDR allows GDR users to ask deeper questions about the origin of datasets, the methods used to collect them, and the findings they help support. Using Retrieval Augmented Generation (RAG), AskGDR can be used to summarize findings spread across dozens of papers and technical reports or to extract relevant information describing a single data field. However, generative AI is experimental. The National Renewable Energy Laboratory (NREL) has been collecting metrics on AskGDR and documenting lessons learned during its deployment. This paper will outline the efficacy and impact of AskGDR through analysis of its use, operating costs, number and types of questions asked, and the quality of answers provided.

15 GEOTHERMAL ENERGY↗

Data Repository for Multi-Objective Urban Observational Strategies: A risk-based framework for expanding flood sensor networks.

These data support the manuscript "Multi-Objective Urban Observational Strategies: A risk-based framework for expanding flood sensor networks." These data are generated to allow water managers to reason about optimal locations to expand a flood observation system from multiple perspectives, specifically focusing on flood hazards, and population exposure to flooding. The data included are a) a shapefile of individual sensor locations b) a shapefile of river reach catchments, c) raster of FEMA flood likelihood layers d) shapefile of population locations and population socioeconomic characteristics. The code is written in R and includes all files necessary to generate the figures for the associated manuscript. Interactive maps of the final calculated maps of hazard, vulnerability, exposure, and risk are also included as html files.

54 ENVIRONMENTAL SCIENCES↗

Consolidated Hydropower Data Repository: Value and Opportunities

Hydropower is one of several types of generating assets that provides energy, capacity, and services to electric power systems. It does so under rubrics and objectives—market driven and regulated, internal and external to asset and fleet owners—that address reliability, cost, price, and, increasingly, flexibility of output. The aggregation of data from multiple hydropower units can provide insights into asset operations and maintenance practices and needs and assist in meeting hydropower objectives. This paper examines the concept and potential benefits of aggregating hydropower asset data—primarily supervisory control and data acquisition (SCADA) information—with examples of insights developed from data aggregated by the Hydropower Research Institute (HRI). Data aggregation as discussed herein, and as implemented by the HRI, extends beyond multiple units in a powerhouse and beyond multiple hydropower facilities in an electric utility fleet or river system. Examples of research and analytics from such aggregated datasets range from unit load dependency analyses to modeling sensor measurements to detect and diagnose anomalies in assets. These examples showed the benefits of utilizing the entire dataset for insights into how the sensor layout of a single unit or set of units compares to the hydropower industry overall. Such insights include whether additional sensors are needed to complete analyses or to make decisions. In addition, utilizing multiple sensors of the same kind within a unit can provide an indication of possible current or upcoming problems with equipment. Although other analyses are possible, their use requires the development of complex models and, potentially, access to types of data that are currently not included with the example dataset used in this study. However, the examples studied herein confirmed the value of data aggregation in the fleet and unit contexts, and the value extends beyond multiple units in a powerhouse and beyond multiple hydropower facilities in an electric utility fleet or river system. The assessments also provided insights into potential extensions to the data aggregation concept that could further add to their value to the hydropower community, and these are included in this document as a set of recommendations.

13 HYDRO ENERGY↗

Recognizability of Demographically Altered Computerized Facial Approximations in an Automated Facial Recognition Context for Potential Application in Unidentified Persons Data Repositories

This study examined the recognizability of demographically altered facial approximations for potential utility in unidentified persons tracking systems. Five computer-generated approximations were generated for each of 26 African male participants using the following demographic parameters: (i) African male (true demographics), (ii) African female, (iii) Caucasian male, (iv) Asian male, and (v) Hispanic male. Overall, 62% of the true demographic facial approximations for the 26 African male participants examined were matched to a corresponding life photo within the top 50 images of a candidate list generated from an automated blind search of an optimally standardized gallery of 6159 photographs. When the African male participants were processed as African females, the identification rate was 50%. In contrast, less congruent identification rates were observed when the African male participants were processed as Caucasian (42%), Asian (35%), and Hispanic (27%) males. The observed results suggest that approximations generated using the opposite sex may be operationally informative if sex is unknown. The performance of approximations generated using alternative ancestry assignments, however, was less congruent with the performance of the true demographic approximation (African male) and may not yield as operationally constructive data as sex-altered approximations.

59 BASIC BIOLOGICAL SCIENCES↗

Oh, My Darling Clementine: A Detailed History and Data Repository of the Los Alamos Plutonium Fast Reactor

The Los Alamos Plutonium Fast Reactor, Clementine, provided the first fast neutron spectrum using a metallic plutonium fuel system and liquid mercury coolant. Clementine is often mentioned in the history of fast, metal-fueled, or metal-cooled designs, with little public information regarding the specifications or timeline. This paper contains a synopsis of each report regarding the development, operation, experimental results, failures, and subsequent disassembly of Clementine to explore this unique reactor within the context of modern advanced reactor development.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Machine Learning-Enhanced Multiphase CFD for Carbon Capture Modeling Run Data

Repository for the data generated as part of the 2023-2024 ALCC project "Machine Learning-Enhanced Multiphase CFD for Carbon Capture Modeling." The data was generated with MFIX-Exa's CFD-DEM model. The problem of interest is gravity driven, particle-laden, gas-solid flow in a triply-periodic domain of length 2048 particle diameters with an aspect ratio of 4. The mean particle concentration ranges from 1% to 40% and the Archimedes number ranges from 18 to 90. The particle-to-fluid density ratio, particle-particle restitution and friction coefficients and domain aspect ratio are held constant at values of 1000, 0.9, 0.25 and 4, respectively. This research used resources of the National Energy Research Scientific Computing Center, a DOE Office of Science User Facility supported by the Office of Science of the U.S. Department of Energy under Contract No. DE-AC02-05CH11231 using NERSC award ALCC-ERCAP0025948.

AMReX↗

The future low-temperature geochemical data-scape as envisioned by the U.S. geochemical community

Data sharing benefits the researcher, the scientific community, and the public by allowing the impact of data to be generalized beyond one project and by making science more transparent. However, many scientific communities have not developed protocols or standards for publishing, citing, and versioning datasets. One community that lags in data management is that of low-temperature geochemistry (LTG). This paper resulted from an initiative from 2018 through 2020 to convene LTG and data scientists in the U.S. to strategize future management of LTG data. Through webinars, a workshop, a preprint, a townhall, and a community survey, the group of U.S. scientists discussed the landscape of data management for LTG – the data-scape. Currently this data-scape includes a “street bazaar” of data repositories. This was deemed appropriate in the same way that LTG scientists publish articles in many journals. The variety of data repositories and journals reflect that LTG scientists target many different scientific questions, produce data with extremely different structures and volumes, and utilize copious and complex metadata. Nonetheless, the group agreed that publication of LTG science must be accompanied by sharing of data in publicly accessible repositories, and, for sample-based data, registration of samples with globally unique persistent identifiers. LTG scientists should use certified data repositories that are either highly structured databases designed for specialized types of data, or unstructured generalized data systems. Recognizing the need for tools to enable search and cross-referencing across the proliferating data repositories, the group proposed that the overall data informatics paradigm in LTG should shift from “build data repository, data will come” to “publish data online, cybertools will find”. Funding agencies could also provide portals for LTG scientists to register funded projects and datasets, and forge approaches that cross national boundaries. Finally, the needed transformation of the LTG data culture requires emphasis in student education on science and management of data.

58 GEOSCIENCES↗

Data from: “Enabling FAIR data in Earth and environmental science with community-centric (meta)data reporting formats”

This dataset contains supplementary information for a manuscript describing the ESS-DIVE (Environmental Systems Science Data Infrastructure for a Virtual Ecosystem) data repository's community data and metadata reporting formats. The purpose of creating the ESS-DIVE reporting formats was to provide guidelines for formatting some of the diverse data types that can be found in the ESS-DIVE repository. The 6 teams of community partners who developed the reporting formats included scientists and engineers from across the Department of Energy National Lab network. Additionally, during the development process, 247 individuals representing 128 institutions provided input on the formats. The primary files in this dataset are 10 data and metadata crosswalk for ESS-DIVE’s reporting formats (all files ending in _crosswalk.csv). The crosswalks compare elements used in each of the reporting formats to other related standards and data resources (e.g., repositories, datasets, data systems). This dataset also contains additional files recommended by ESS-DIVE’s file-level metadata reporting format. Each data file has an associated dictionary (files ending in _dd.csv) which provide a brief description of each standard or data resource consulted in the data reporting format development process. The flmd.csv file describes each file contained within the dataset.

54 ENVIRONMENTAL SCIENCES↗

Opening doors to physical sample tracking and attribution in Earth and environmental sciences

Physical samples and their associated data and metadata underpin scientific discoveries across disciplines and can enable new science when appropriately archived. However, there are significant gaps in current practices and infrastructure that prevent accurate provenance tracking, reproducibility, and attribution. For most samples, descriptive metadata are often sparse, inaccessible, or absent. Samples and associated data and metadata may also be scattered across numerous physical collections, data repositories, laboratories, data files, and papers with no clear linkage or provenance tracking as new information is generated over time. The Earth Science Information Partners (ESIP) Physical Samples Curation Cluster has therefore developed guidance for scientific authors on ‘Publishing Open Research Using Physical Samples.’ This involved synthesizing existing practices, gathering community feedback, and assessing real-world examples. We identified improvements needed to enable authors to efficiently cite and link Earth science samples and related data, and track their use. Our goal is to help improve discoverability, interoperability, and reuse of physical samples, and associated data and metadata. Though primarily focused on the needs of Earth and environmental sciences, these guidelines are broadly applicable.

58 GEOSCIENCES↗

MultiProBE Software Platform

MultiProBE Software Platform is cloud based multi-omic data repository and data processing compute platform.

Wright, Devin↗

Data for publication: "A fresh take: Seasonal changes in terrestrial freshwater inputs impact salt marsh hydrology and vegetation dynamics"

This data repository contains data associated with the manuscript "A fresh take: Seasonal changes in terrestrial freshwater inputs impact salt marsh hydrology and vegetation dynamics". This study was conducted at the Elkhorn Slough National Estuarine Research Reserve in Watsonville, California from October 2019 - June 2022. We sought to understand the role of shallow freshwater inputs from adjacent uplands on salt marsh hydrologic behavior and vegetation productivity. This dataset contains CSV files of the following: daily salt marsh subsurface water level and pore water conductivity, monthly vegetation survey measurements, soil core data. Estuary surface water level, conductivity, and local precipitation were downloaded from the National Estuarine Research Reserve System (Centralized Data Management Office) at https://cdmo.baruch.sc.edu/.

54 ENVIRONMENTAL SCIENCES↗

Integrated GW Farm ABM

This Data Repository includes data used for the integrated groundwater- farm ABM model, raw model output from scenario ensemble, and processed outputs that isolate the groundwater storage depletion outcomes for the 35,000 farm cells. Model Inputs: Farm ABM Inputs: This folder contains the input data used by the integrated groundwater - farm ABM modelling script (Python file) used for the high performance computing (HPC) experiments. The sub-folder "data inputs" contains all of the farm attribute data, while the three files in the folder have the hydrogeological data lookup table (NLDAS Cost Curve Attributes.csv), a lookup table (Theis well function table.csv) for the groundwater cost curve function, and the farm indexes and corresponding NLDAS ids for all of the cells run in this experiment (nldas farms subset final.csv). NLDAS Cost curve hydrogeological data: Hydrogeological data aggregated to 1/8 degree resolution and aligned with the NLDAS grid. Parameters include: water depth below ground surface [meters], subsurface porosity [unitless], aquifer depth from ground surface to aquifer bottom [meters], annual average recharge (USGS: mm, Doll: meters), and three different hydraulic conductivity (K) values (meters/day). The three K values represent the mean value from Gleeson et al. (2018), one standard deviation above the mean from Gleeson et al. (2018), and the de Graaf et al. 2020 modifications to certain lithologies. Additional information about these datasets and their processing are documented in the supplement to Yoon et al. 2025 (in review). Output: Raw outputs: This folder contains a .zip file that has model outputs for the entire scenario ensemble. There is one csv for each farm id, using the format "farm farmid cases.csv". The relationship between the farm id and NLDAS id is defined by the "nldas farms subset final.csv" located in the Farm ABM Inputs folder. Each csv has 625 rows, corresponding to 625 combinations of different scenario parameter values. Each row (scenario) represents the outcome of a 100 year simulation. Columns define scenario settings and summary statistics for each scenario. The first four columns define the scenario settings: "hydro ratio," "econ ratio," "K scenario," and "gamma scenario." The hydro and econ ratios are values passed to the modeling script that influence multipliers for other model parameters, as documented in the supplement to Yoon et al. 2025 (in review). The gamma multiplier is a coefficient multiplier applied to the baseline gamma values (values below 1 represent lower unobserved costs compared to baseline, values above 1 represent higher costs). The K scenario names represent K values of: "low": 0.5 m/d, "int 1": 2.5 m/d, "int 2": 10 m/d, "high": 50 m/d, and "gleeson": mean Gleeson K value. "Perc vol depleted" is the fraction of groundwater depleted at the end of the 100 simulation. Processed Output: Derived depletion outcomes from raw outputs: All of the individual csv files from the Raw outputs were aggregated into a single file that has the scenario settings and fraction depletion "Perc vol depleted" for every farm cell, for every scenario. The other two files define relationships between the farm id, NLDAS id, and local and major aquifer units, used for aquifer-level depletion analysis.

Agent based modeling↗

MSD CoP Webinar: "Advances in MSD-LIVE to Support the MSD Community of Practice"

Context: This webinar was hosted by the MultiSector Dynamics Community of Practice (MSD CoP; https://multisectordynamics.org). Advances in MSD-LIVE to Support the MSD Community of Practice Presenters: Casey Burleyson and Zoe Guillen (Pacific Northwest National Laboratory) Abstract: The MultiSector Dynamics Living, Intuitive, Value-adding, Environment (MSD-LIVE; msdlive.org) is a cloud-based data management system and advanced computing platform that enables MSD researchers to document and archive their data, run their models and analysis tools, and share their data, software, and workflows within the MSD Community of Practice. Recently, several high-profile datasets have attracted many new users to MSD-LIVE. This webinar has two goals: 1) To refamiliarize the MSD community and new users with the components of the platform (e.g., the data repository, model training notebooks, and data dashboards) and to highlight examples of how these components are advancing MSD science and 2) To demonstrate new features in v3 of the platform, released in late 2025. The main new feature in v3 is the ability to interactively explore data in MSD-LIVE without downloading it. MSD-LIVE users can now click a button in our data repository and launch a blank Jupyter notebook with access to the underlying data on AWS. Users can use the notebook to write analysis, visualization, or subsetting routines that process the data directly on the AWS cloud. We also added a GitHub integration feature that allows users to share analysis or visualization code they develop with the community of MSD-LIVE users. The webinar will wrap up with a look at what's coming next for MSD-LIVE in 2026. Moderator: Patrick M. Reed (MSD CoP Facilitation Team) This webinar was held on: May 12th, 2026 from 1-2 PM EST.

Open Science↗