Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data repository”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Network Analysis of Academic Medical Center Websites in the United States

Healthcare resources are published annually in repositories such as the AHA Annual Survey Database TM . However, these data repositories are created via manual surveying techniques which are cumbersome in collection and not updated as frequently as website information of the respective hospital systems represented. Also, this resource is not widely available to patients in an easy-to-use format. Network analysis techniques have the potential to create topological maps which serve to aid in pathfinding for patients in their search for healthcare services. This study explores the topological structure of forty United States academic health center websites. Network analysis is utilized to analyze and visualize 48,686 webpages. Several elements of network structure are examined including basic network properties, and centrality measures distributions. The Louvain community detection algorithm is used to examine the extent to which these techniques allow identification of healthcare resources within networks. The results indicate that websites with related healthcare services tend to form observable clusters useful in mapping key resources within a hospital system.

97 MATHEMATICS AND COMPUTING↗

Cloud-Feedback Model Intercomparison Project: Tier 2 Simulations (Final Report)

The University of Miami (Subcontractor)’s Research Scientist James Benedict (with oversight by PI Amy Clement) was tasked with producing global climate model simulations as part of the Cloud-Feedback Model Intercomparison Project (CFMIP) Tier 2 protocol, managing the output of these simulations, providing assistance (as requested) to staff at Lawrence Livermore National Lab (LLNL) regarding model output and setup, and meeting with LLNL scientists and team members to coordinate work on the project. The simulations represent an important contribution to the CFMIP data repository and will advance understanding of a wide range of critical cloud, circulation, and precipitation responses to climate change. The Subcontractor completed all proposed simulations, formatted the model output to be compliant with CFMIP protocols, and published the model data to the CFMIP repository. When requested, the Subcontractor provided assistance to LLNL for model output and configuration queries. Benedict and/or Clement also met in person or virtually 1-2 times per year with project scientists to coordinate work and review results.

54 ENVIRONMENTAL SCIENCES↗

Data and code repository for "How do the weather regimes drive wind speed and power production at the sub-seasonal to seasonal timescales over the CONUS?"

There has been an increasing need for forecasting power generation at the sub-seasonal to seasonal (S2S) timescales to support the operation, management, and planning of the wind-energy system. At the S2S timescales, atmospheric variability is largely related to recurrent and persistent weather patterns, referred to as weather regimes (WRs). In the study "How do the weather regimes drive wind speed and power production at the sub-seasonal to seasonal timescales over the CONUS?", we identify four WRs that influence wind resources over North America using a self-organizing map (SOM) algorithm. These WRs are responsible for large-scale wind and power production anomalies over the CONUS at the S2S timescales. The WR-based reconstruction explains up to 50% of the monthly variance of power production over the western United States, and the explanatory power generally increases with the increase of timescales. The identified relationship between WRs and power production reveals the potential and limitations of the regional WR-based wind resource assessment over different regions of the CONUS across multiple timescales. This repository includes all the data and codes we use for analyses in this study. Users may use them to reproduce the results of this study on their end.

17 WIND ENERGY↗

Projections of Hourly Meteorology by County Based on the IM3/HyperFACETS Thermodynamic Global Warming (TGW) Simulations

This dataset contains 40 years (1980-2019) of historical hourly meteorology and 80 years (2020-2099) of projected hourly meteorology for each county in the conterminous United States. Details about the scenarios and variables included in this dataset are in the readme.pdf file. This dataset is derived from the IM3/HyperFACETS Thermodynamic Global Warming (TGW) simulations (https://doi.org/10.57931/1885756). More details on the TGW approach can be found at: https://tgw-data.msdlive.org/. If you use this dataset please also cite the raw TGW dataset (Jones, A. D., Rastogi, D., Vahmani, P., Stansfield, A., Reed, K., Thurber, T., Ullrich, P., & Rice, J. S. (2022). IM3/HyperFACETS Thermodynamic Global Warming (TGW) Simulation Datasets (v1.0.0) [Data set]. MSD-LIVE Data Repository. https://doi.org/10.57931/1885756). The four future climate scenarios in the TGW data are: rcp45cooler, rcp45hotter, rcp85cooler, and rcp85hotter. For the historical period and each of the four future scenarios the dataset has hourly estimates of the spatial-average of six meteorological variables for each county: Temperature, specific humidity, shortwave radiation, longwave radiation, and the U (east-west) and V (north-south) components of the wind speed. Times for each file are in the filename and all times are in Coordinated Universal Time (UTC). Counties are identified by their Federal Information Processing Standard (FIPS) code. The mapping between counties and FIPS codes is provided in the state_and_county_fips_codes.csv file. The code to go from the raw TGW data to county-level projections is available at: https://github.com/IMMM-SFA/im3components/tree/main/im3components/wrf_to_tell.

County↗

FAIR Interfaces for Geospatial Scientific Data Searches

Several factors must be considered in designing a highly accurate, reliable, scalable, and user-friendly geospatial data search interfaces. This paper examines four critical questions that ought to be considered during design phase: (1) Is the search interface or API that provides the search capability useable by both humans and machines? (2) Are the results consistent and reliable? (3) Is the output response format free to use, community-defined, and non-propriety? (4) Does the API clearly state the usage clauses? This paper discusses how certain data repositories at the US Department of Energy's Oak Ridge National Laboratory apply FAIR data principles to enable geospatial searches and address the above-mentioned questions.

Devarakonda, Ranjeet↗

Retrospective on Recent DOE-Funded Studies Concerning the Extraction of Rare Earth Elements & Lithium from Geothermal Brines (Final Report)

Rare earth elements (REE) and lithium are non-toxic metals that are considered critical materials due to their use in electronics, magnets, batteries, and a wide variety of industrial processes important for the economy and military preparedness. Demand for REE and lithium is increasing and these critical materials are imported, so identifying and exploiting domestic sources of REE and lithium is a national priority. The U.S. Department of Energy (DOE) Geothermal Technologies Office (GTO) has been in the forefront of sponsoring research investigating the potential recovery of REE, lithium, and other critical minerals from geothermal brines. It has been proposed that the future of geothermal energy should include “hybrid systems” that combine electricity generation with other revenue-generating activities, such as recovery of valuable and critical minerals, including REE and lithium. Two recent GTO funding opportunities have focused on the recovery of REE and other valuable minerals from geothermal brines. The research supported by the GTO’s mineral recovery program is focused on three areas: resource characterization, technology for the extraction of REE, and technology for the extraction of lithium (Tables 1 and 2). This report is a retrospective study examining the outcome of GTO’s two recent mineral recovery programs (DE-FOA-0001016 in FY 2014 and DE-FOA-0001376 in FY 2016). In this report, the knowledge, technology, and techniques that were developed by researchers funded by GTO are summarized and discussed. Four projects were funded to assess the concentrations and amounts of REE found in geothermal brines and oil field produced waters. The GTO-funded studies compiled publically available data on REE concentrations from brines and produced water from all over the USA. In addition, new samples were collected and characterized from major geothermal and hydrocarbon basins in the Western USA. The studies examined the relationship between lithology and REE concentrations and developed models examining the influence of geology on REE concentrations in produced brines. It was determined that REE are frequently found at higher concentrations in oil field produced water than geothermal brines, but that some geothermal areas had significant REE resources. Significant reservoirs of REE were identified in the Western USA. In some cases, concentrations of REE were more than 1000 times the concentrations found in seawater. Collectively, these studies represent a comprehensive picture of REE resources associated with geothermal and hydrocarbon systems in the USA. The studies did not examine lithium resources, but in some cases, lithium concentration data was collected. Data from these studies are housed in the Geothermal Data Repository (GDR) and represent a significant information resource and it is recommended that these data be further analyzed in a future study. Eight projects were funded to develop new technology for REE extraction from geothermal fluids. These projects investigated sorption as an approach for removal and recovery of REE from geothermal brines. The projects investigated cutting-edge technology for selective sorption of ions from complex solutions, including the application of metal-organic frameworks and biosorbent proteins. The REE sorption studies tested different combinations of metal-binding ligands and solid supports. The most promising metal-binding ligands for REE included phosphonic acid, thiol, and carboxylic acid functional groups. Ligands were attached or incorporated into a wide variety of solid supports. In most cases, attachment was via covalent bonding to organic resins, polymers, or silica-based supports. Most of the REE projects were conducted at a low technology readiness level (TRL) and showed promise, but direct comparison between technologies was not possible based on the available information. It is recommended that testing and reporting be standardized to the extent possible to facilitate comparisons between technologies. Two projects were directed at novel lithium extraction technology. Both projects investigated the use of inorganic sorbents, including manganese oxides. One study also examined the use of metal- ion imprinted polymers as selective ion-exchange resins for the separation of lithium and manganese from brines. Both approaches showed promise for the selective extraction of lithium from brines, including potentially geothermal brines. Results from these GTO studies indicated that selective REE and lithium extraction is possible, but interference from co-occurring solutes, such as calcium, magnesium, or heavy metals, will interfere with process efficiency and negatively impact process economics. Techno-economic analysis conducted as part of the resource and technology studies suggest extraction of REE from geothermal brines is unlikely to be economically viable, especially since non-geothermal produced waters frequently have higher REE concentrations. It is recommended that benchmarks for techno-economic analysis be established to the extent possible for future studies, to facilitate direct comparison of various technologies. Based on the collective results of this program, it appears that hybrid geothermal power would benefit more from recovery of lithium and other metals, rather than REE. It is recommended that future studies be conducted at a higher- TRL and that sorbents be tested against actual geothermal fluid samples. Prior higher-TRL efforts to extract metals from geothermal brines should be further evaluated for lessons learned.

36 MATERIALS SCIENCE↗

University of Kentucky measurements of wind, temperature, pressure and humidity in support of LAPSE-RATE using multisite fixed-wing and rotorcraft unmanned aerial systems

In July 2018, unmanned aerial systems (UASs) were deployed to measure the properties of the lower atmosphere within the San Luis Valley, an elevated valley in Colorado, USA, as part of the Lower Atmospheric Profiling Studies at Elevation – a Remotely-piloted Aircraft Team Experiment (LAPSE-RATE). Measurement objectives included detailing boundary layer transition, canyon cold-air drainage and convection initiation within the valley. Details of the contribution to LAPSE-RATE made by the University of Kentucky are provided here, which include measurements by seven different fixed-wing and rotorcraft UASs totaling over 178 flights with validated data. The data from these coordinated UAS flights consist of thermodynamic and kinematic variables (air temperature, humidity, pressure, wind speed and direction) and include vertical profiles up to 900 m above the ground level and horizontal transects up to 1500 m in length. These measurements have been quality controlled and are openly available in the Zenodo LAPSE-RATE community data repository (https://zenodo.org/communities/lapse-rate/, last access: 23 July 2020), with the University of Kentucky data available at https://doi.org/10.5281/zenodo.3701845 (Bailey et al., 2020).

54 ENVIRONMENTAL SCIENCES↗

Estimating Critical Customer Outages Resulting from Extreme Hurricanes

US power outage data has been collected by organizations such as Oak Ridge National Laboratory (ORNL) through Environment for Analysis Geo-Located Energy Infrastructure (EAGLE-I: freely available) and poweroutage.us (commercial data: available to purchase). However, these sources do not provide information specific to outages of critical customers. Critical customers include entities, facilities, and individuals whose continuous access to electricity is essential for public safety, emergency response, disaster recovery, the well-being of vulnerable populations, public safety and order, and public utilities such as natural gas, communications, water and sanitation. Identification and geolocation of critical customers is crucial for understanding and addressing the effects of power outages on essential services and ensuring that necessary measures are taken to maintain their operations during power disruptions. This work is a first step towards estimating the occurrences of critical customer outages and developing a critical customer power outage data repository. This work estimates outage incidents of critical customers through spatiotemporal mapping of power outage data, weather data, building data, and critical infrastructure network data. Our results show that critical customer effects vary across different counties. We provide appropriate mathematical explanations and simplifications to define and systematize the proposed approach.

Bhusal, Narayan [ORNL] (ORCID:0000000222752145)↗

Cyote Insights

CyOTE Insights leverages React, Vite, Typescript, Tailwind, and Daisy UI for the Graphical User Interface. It was designed in a particular style with a dark mode and a light mode. All code is broken down into components and reusable wrapper components for efficiency. All data is stored in Deep Lynx as a central data repository using an ontology based schema. The application serves as a main endpoint for the data in the COREII and CyOTE programs. The main purpose of the application is to display historical attack data in the Operational Technology space. At the time of this writing, it supports 27 historical attack reports compiled from OSINT sources. All of the data is publicly available, but what this application offers is the ability to see many years worth of publications in a detailed dashboard. It will also support future reports that are written using the other applications in the COREII program.

Pluth, AdamJ [Idaho National Laboratory (INL), Ida↗

A database of hourly wind speed and modeled generation for US wind plants based on three meteorological models

Abstract In 2022, wind generation accounted for ~10% of total electricity generation in the United States. As wind energy accounts for a greater portion of total energy, understanding geographic and temporal variation in wind generation is key to many planning, operational, and research questions. However, in-situ observations of wind speed are expensive to make and rarely shared publicly. Meteorological models are commonly used to estimate wind speeds, but vary in quality and are often challenging to access and interpret. The Plant-Level US multi-model WIND and generation (PLUSWIND) data repository helps to address these challenges. PLUSWIND provides wind speeds and estimated generation on an hourly basis at almost all wind plants across the contiguous United States from 2018–2021. The repository contains wind speeds and generation based on three different meteorological models: ERA5, MERRA2, and HRRR. Data are publicly accessible in simple csv files. Modeled generation is compared to regional and plant records, which highlights model biases and errors and how they differ by model, across regions, and across time frames.

17 WIND ENERGY↗

The Data Foundry: Secure Collaboration for the Geothermal Industry: Preprint

The Data Foundry provides secure, cloud-based storage and universal access to digital information, enabling the greater geothermal industry to collaborate seamlessly with the Department of Energy (DOE), national labs, universities, and private organizations. Originally developed to support the EGS Collab project, the Data Foundry has been expanded to support FORGE, EDGE, and other DOE-funded projects, some collaborative, some private, by providing each project with a secure space and the ability to fine-tune individual access controls. In response to user feedback, it now also features improved integration with DOE’s Geothermal Data Repository (GDR), to provide a clear and convenient pathway from collaboration to publication, and to register collaborative data projects with well-known data registries like Data.gov and the National Geothermal Data System (NGDS). This paper will explore how recent concerns raised by data-centric, proprietary projects have informed development on the Data Foundry and highlight improvements designed to streamline workflows, improve access control, and promote the timely dissemination of information to the geothermal industry.

Data Foundry↗

Machine Learning (ML) Classifier to Assist Metadata Creation

The Atmospheric Radiation Measurement (ARM) Data Center is responsible for the timely collection, archival, and curation of science data products. These products are freely available through an online data repository. Metadata creation is paramount for scientific users to find and access over seven petabytes of atmospheric science data. The hierarchical metadata structure allows users to search for information at both broad and narrow levels. This project aims to leverage 30 years’ worth of manually created metadata to enable machine predictions of broad-term classifications from narrow-term descriptions. These classification predictions would assist metadata coordinators with their term selections. This paper discusses the cleaning and preprocessing of the training data, the pipeline developed to determine the best model for this task, and the creation of an API metadata classifier for ARM measurement metadata. Our results show that the Linear Support Vector Classification (LinearSVC) algorithm, along with the Term Frequency – Inverse Document Frequency (TF-IDF) vectorizer, is well-suited for our multi-class classification task. Lengthier input training data led to better results, and artificial balancing was unnecessary for this particular use case. This predictive classifier enhances efficiency in metadata creation, as well as supports greater consistency and accuracy in metadata tagging.

Collier, Hannah [ORNL] (ORCID:0000000341284292)↗

Utilizing Ontology Structures To Curate the DOE-NETL Carbon Storage Open Database

The specialized ontology for the Carbon Storage Open Database will enable more rapid assignment of appropriate symbology standards for visualization improvements, optimize topical and spatial tagging within keywords, and improve flexibility for utilization in existing data repositories such as EDX. This effort also aims to establish a foundation for utilization of ontologies for organization of other data related to geologic carbon storage in the future.

Martin, Abigail↗

Efforts to enhance reproducibility in a human performance research project

Background: Ensuring the validity of results from funded programs is a critical concern for agencies that sponsor biological research. In recent years, the open science movement has sought to promote reproducibility by encouraging sharing not only of finished manuscripts but also of data and code supporting their findings. While these innovations have lent support to third-party efforts to replicate calculations underlying key results in the scientific literature, fields of inquiry where privacy considerations or other sensitivities preclude the broad distribution of raw data or analysis may require a more targeted approach to promote the quality of research output. Methods: We describe efforts oriented toward this goal that were implemented in one human performance research program, Measuring Biological Aptitude, organized by the Defense Advanced Research Project Agency's Biological Technologies Office. Our team implemented a four-pronged independent verification and validation (IV&V) strategy including 1) a centralized data storage and exchange platform, 2) quality assurance and quality control (QA/QC) of data collection, 3) test and evaluation of performer models, and 4) an archival software and data repository. Results: Our IV&V plan was carried out with assistance from both the funding agency and participating teams of researchers. QA/QC of data acquisition aided in process improvement and the flagging of experimental errors. Holdout validation set tests provided an independent gauge of model performance. Conclusions: In circumstances that do not support a fully open approach to scientific criticism, standing up independent teams to cross-check and validate the results generated by primary investigators can be an important tool to promote reproducibility of results.

59 BASIC BIOLOGICAL SCIENCES↗

High Energy Physics Network Requirements Review: Two-Year Update

The Energy Sciences Network (ESnet) is the high-performance network user facility for the US Department of Energy (DOE) Office of Science (SC) and delivers highly reliable data transport capabilities optimized for the requirements of data-intensive science. In essence, ESnet is the circulatory system that enables the DOE science mission by connecting all its laboratories and facilities in the US and abroad. ESnet is funded and stewarded by the Advanced Scientific Computing Research (ASCR) program and managed and operated by the Scientific Networking Division at Lawrence Berkeley National Laboratory (LBNL). ESnet is widely regarded as a global leader in the research and education networking community. ESnet interconnects DOE national laboratories, user facilities, and major experiments so that scientists can use remote instruments and computing resources as well as share data with collaborators, transfer large datasets, and access distributed data repositories. ESnet is specifically built to provide a range of network services tailored to meet the unique requirements of the DOE’s data-intensive science. In July 2023, the Energy Sciences Network (ESnet) and the High Energy Physics program (HEP) of the DOE SC organized an interim ESnet requirements review of HEP-supported activities, to follow up on the work started during the 2020 HEP Network Requirements Review. Preparation for these events included checking back with the key stakeholders: program and facility management, research groups, and technology providers. Each stakeholder group was asked to prepare updates to their previously submitted case study documents, so that ESnet could update the understanding of any changes to the current, near-term, and long-term status, expectations, and processes that will support the science activities of the program.

97 MATHEMATICS AND COMPUTING↗

Guiding the choice of informatics software and tools for lipidomics research applications

Progress in mass spectrometry lipidomics has led to a rapid proliferation of studies across biology and biomedicine. These generate extremely large raw datasets requiring sophisticated solutions to support automated data processing. To address this, numerous software tools have been developed and tailored for specific tasks. However, for researchers, deciding which approach best suits their application relies on ad hoc testing, which is inefficient and time consuming. Here we first review the data processing pipeline, summarizing the scope of available tools. Next, to support researchers, LIPID MAPS provides an interactive online portal listing open-access tools with a graphical user interface. This guides users towards appropriate solutions within major areas in data processing, including (1) lipid-oriented databases, (2) mass spectrometry data repositories, (3) analysis of targeted lipidomics datasets, (4) lipid identification and (5) quantification from untargeted lipidomics datasets, (6) statistical analysis and visualization, and (7) data integration solutions. Detailed descriptions of functions and requirements are provided to guide customized data analysis workflows.

59 BASIC BIOLOGICAL SCIENCES↗

Oak Ridge National Laboratory Pilot Demonstration of an Attestation and Anomaly Detection Framework using Distributed Ledger Technology for Power Grid Infrastructure

This report summarizes the design and pilot demonstration of a framework called Grid Guard that was created to provide increased data and device trustworthiness to electric grid devices by leveraging distributed ledger technology (DLT), specifically blockchain. Grid Guard contains a combination of core cryptographic methods such as the secure hash algorithm (SHA), and asymmetric cryptography, private permissioned blockchain, baselining configuration data, consensus algorithm (Raft) and the Hyperledger Fabric (HLF) framework. The system implements a low energy, fast, and robust enhancement to system trustworthiness within and across electric grid systems such as substations, control centers and metering infrastructures. Blockchain is a distributed database structured that provides a practically unalterable (immutable) timeline of stored transactions. By relying on hashing and the Raft consensus algorithm, if an entity tries to illegitimately alter a record at one instance of the database the other ledger nodes are not altered. They work to cross-reference each other and easily locate any incorrectly added data and remove it. The bulk raw data is stored in an off-chain storage (outside of the blockchain ledger) and a hash of this baseline data is stored in the Blockchain ledger via hashing windows of time-series and configuration data, after aggregation and filtering. The bulk off-chain data repository is then considered to be trust-anchored using the hashes stored in the blockchain. To secure the electric grid testbed devices and data, device configuration baselines were compared to those baselines that had been previously stored in the ledger. Statistical baselines for device configurations, network communication patterns, and high-speed sensor data are calculated and then stored off-chain and hashes stored in the ledger. Measurements such as three-phase voltage and current, frequency, breaker status, protection scheme settings, network configuration settings (and other device configuration artifacts) and network traffic features (packet interarrival times) are compared every minute or other selected time windows. During phase 1 of the Grid Guard DLT project different DLT technologies were studies, and an assessment was performed on DLT technology vulnerabilities, uses, and key characteristics. DLT consensus protocols were studies (e.g., RAFT, named after Reliable, Replicated, Redundant, And Fault-Tolerant). Also, cryptography, public, private and permissioned or permissionless systems were assessed. Grid Guard implements a permissioned private DLT. Consensus algorithm selection and choice of DLT implementation depended heavily on the use-case. For this use-case, parameters were selected to measure performance and existing tools for assessment. Benchmarking was performed theoretically and practically. During phase 2 hashed transactions/blocks were inserted into the ledger every second. During phase 2 of the Grid Guard DLT project, a prototype framework was developed and demonstrated for attestation of critical substation devices and data using precision timing systems that use PTP and IRIG-B protocols) on a testbed of operational devices that emulated a distribution substation, control center, and power metering infrastructure using real Operational Technology (OT). The testbed includes OT devices such as protective relays, human machine interfaces (HMI), and power meters. To determine when to collect and compare system and network baselines, an initial examination of an anomaly detection capability to identify malicious manipulation of data streams was conducted. The resulting anomaly detection was demonstrated in a set of experiments and leveraged to trigger device artifact attestation checks. Attestation checks occur against device configuration baselines when compared with the immutable blockchain-stored baselines, which provided a cryptographically supported means by which to store baselines. The electrical substation-grid testbed was created to test the Grid Guard framework. The testbed emulates the operations of a portion of a power grid and SCADA systems as closely as possible. The testbed integrates real protocols, mainly IEC 61850 standard protocols, such as the Sampled Value (SV) and the GOOSE protocols. The testbed also supports DNP3 and other layer 2 and layer 3 protocols such as Telnet, SSH, SFTP/FTP and other proprietary protocols needed to connect to industrial control system equipment. The testbed emulates real power conditions using the OpalRT hardware-in-the-loop (HIL) device which can create fault situations that cannot be easily tested on real systems. The electrical substation-grid testbed was created using real measurement, communication, and protection devices that electrical utilities commonly use.

24 POWER TRANSMISSION AND DISTRIBUTION↗