Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Sources”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Development of National New Construction Weighting Factors for the Commercial Building Prototype Analyses (2008-2022)

The U.S. Department of Energy (DOE) tasked Pacific Northwest National Laboratory (PNNL) with updating commercial building construction weights for the purpose of estimating national and state-by-state energy savings impacts of changes made to various commercial energy codes and standards. A similar activity was last completed by PNNL in 2020 using disaggregate construction volume data acquired from the Dodge Data & Analytics database (formerly McGraw Hill) for the years 2003-2018 (Lei et al, 2020). As time passes, changes in economic and social demand reshape construction volume trends. For the current update, PNNL reviewed the same data source with the latest construction data for the years 2008-2022. For commercial building analyses, PNNL typically uses a suite of 16 prototype buildings simulated in the 19 ASHRAE climate zones with 16 of them present in the United States. The 2008-2022 commercial building weighting factors were derived using the same approach employed to develop the 2003-2018 set (Lei et al, 2020). Applying the construction volume data from the database to the prototypes and climate zones resulted in the following new construction area-based weighting factors. Table ES.1 shows the weighting factors including all building categories found in the database, and Table ES.2 shows the weighting factors normalized to include only buildings represented by the 16 prototypes. Section 3.0 also includes national- and state-level weighting factors by area and building count.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

A Play-Based Exploration of CO 2 Storage in the Illinois Basin (Final Technical Report)

This report documents the results of "A Play-Based Exploration of CO₂ Storage in the Illinois Basin" (DE-FE0032366), a project funded by the U.S. Department of Energy Office of Fossil Energy and Carbon Management and conducted by the Illinois State Geological Survey (ISGS) at the University of Illinois Urbana-Champaign in partnership with Visage Energy. The project adapted play-based exploration (PBE), a systematic basin-scale evaluation methodology from the petroleum industry, to screen areas of Illinois for commercial geologic carbon storage (GCS) in Cambro-Ordovician strata. The traditional play concept was expanded to encompass three play element groups, subsurface geologic factors, surface features, and societal factors, yielding 24 play elements with defensible suitability criteria applied through a five-tier classification scheme. An integrated geospatial database was assembled from ISGS, MGSC, MRCI, NATCARB, and public data sources, supported by significant data-improvement work including correction of legacy well locations, digitization of more than 1,300 well construction records using the DOE CATALOG team's OGRRE tool, compilation of a statewide 2D seismic database, and production of a refined fault and fold geodatabase.

58 GEOSCIENCES↗

KS4A-Omics1.0_FspDS682

Soil fungi facilitate the translocation of inorganic nutrients from soil minerals to other microorganisms and plants. This ability is particularly advantageous in impoverished soils, because fungal mycelial networks can bridge otherwise spatially disconnected and inaccessible nutrient hotspots. However, the molecular mechanisms underlying fungal mineral weathering and transport through soil remains poorly understood. Here, we addressed this knowledge gap by directly visualizing nutrient acquisition and transport through fungal hyphae in a mineral doped soil micromodel using a multimodal imaging approach. Here, we observed how a representative of common saprotrophic soil fungi, Fusarium sp. DS 682, exhibited a mechanosensory response (thigmotropism) around obstacles and through pore spaces (~12 μm) in the presence of minerals.This study establishes the significance of fungal biology and nutrient translocation mechanisms in maintaining fungal growth under water and nutrient limitations in a soil-like microenvironment, using a high-throughput multi-omic analysis approach. Data package KS4A-Omics.1.0_FspDS682 (Publication: Fungal Mineral Weathering Mechanisms Revealed Through Direct Molecular Visualization) contents reported here are the first version (1.0) and contain pre- and post-processed data using high throughput data capture technologies for multi-omic analysis and integration, this data package contains raw and post-processed experimental data for X-Ray Absorption Near Edge Structure Spectroscopy (XANES/XRF), Optical Microscopy, Proteomics, Scanning Electron Microscope (SEM), Time-of-Flight Secondary Ion Mass Spectroscopy (ToF-SIMS), X-Ray Diffraction Spectroscopy (XRD) files, and X- Ray Photoelectron Spectroscopy (XPS) using EMSL capabilities. This data package DOI contains a comprehensive collection of high-throughput multi-omics data and process metadata catalog. Support files include additional data download contents “Read Me” with dataset descriptor information and data source method application ontologies (see data dictionary section). Reported data download content is structured for compliance with reported guidelines provided by community standard initiatives and publisher stakeholder policies supporting FAIR data principles.

47 OTHER INSTRUMENTATION↗

Latency Analysis of the Nexus Digital Twin Framework

Real-time digital catalogs are increasingly relied upon to track metadata and connect disparate data sources for cloud-based data integration efforts. One such tool, Deeplynx Nexus is supporting real-time digital twin efforts through event-driven data integration and time-series queries. Nexus’s usefulness for these applications depends critically on how quickly individual records can be uploaded and downloaded, since delays directly affect the responsiveness of any system built on top of it. However, the actual latency a user should expect from Nexus has not been systematically measured before, particularly for the small, frequent transactions typical of live sensor feeds. Here we show that single-record round-trip latency is 61.1 ms on a local Nexus instance and 391.7 ms on the hosted production infrastructure, a roughly 6.4x difference driven primarily by fixed per-request overhead rather than data volume. This overhead dominates at small scale: comparing single-record and ten-record trials suggests approximately 56 ms of each single-record request is fixed connection and authentication cost rather than data-transfer time, meaning batching even a handful of records is substantially more efficient than transmitting them individually. At large batch sizes, this pattern reverses for uploads, which converge to near parity between local and hosted environments by 25,000-50,000 records, while download latency remains persistently 5.7-6.4x slower on hosted infrastructure even at scale. These results suggest that Nexus deployments intended for real-time digital twin applications should prioritize record batching over single-record transactions, and that download-path optimization on hosted infrastructure offers the largest remaining opportunity to reduce latency at scale. We anticipate these baseline measurements will serve as a reference point for future digital twin projects evaluating whether Nexus’s latency profile meets their real-time requirements, and as a benchmark for tracking the effect of future infrastructure or API changes.

99 - GENERAL AND MISCELLANEOUS↗

Testing SOAR tools in use

Investigations within Security Operation Centers (SOCs) are tedious as they rely on manual efforts to query diverse data sources, overlay related logs, correlate the data into information, and then document results in a ticketing system. Security Orchestration, Automation, and Response (SOAR) tools are a relatively new technology that promise, with appropriate configuration, to collect, filter, and display needed diverse information; automate many of the common tasks that unnecessarily require SOC analysts’ time; facilitate SOC collaboration; and, in doing so, improve both efficiency and consistency of SOCs. There has been no prior research to test SOAR tools in practice; hence, understanding and evaluation of their effect is nascent and needed. Here, in this paper, we design and administer the first hands-on user study of SOAR tools, involving 24 participants and six commercial SOAR tools. Our contributions include the experimental design, itemizing six characteristics of SOAR tools, and a methodology for testing them. We describe configuration of a cyber range test environment, including network, user, and threat emulation; a full SOC tool suite; and creation of artifacts allowing multiple representative investigation scenarios to permit testing. We present the first research results on SOAR tools. Concisely, our findings are that: per-SOC SOAR configuration is extremely important; SOAR tools increase efficiency and reduce context switching, although with potentially decreased ticketing accuracy/completeness; user preference is slightly negatively correlated with their performance with the tool; internet dependence varies widely among SOAR tools; and balance of automation with assisting decision making is preferred by senior participants. We deliver a public user- and tool-anonymized and -obfuscated version of the data.

97 MATHEMATICS AND COMPUTING↗

Evaluation and Intercomparison of Small Uncrewed Aircraft Systems Used for Atmospheric Research

Abstract Small uncrewed aircraft systems (sUAS) are regularly being used to conduct atmospheric research and are starting to be used as a data source for informing weather models through data assimilation. However, only a limited number of studies have been conducted to evaluate the performance of these systems and assess their ability to replicate measurements from more traditional sensors such as radiosondes and towers. In the current work, we use data collected in central Oklahoma over a 2-week period to offer insight into the performance of five different sUAS platforms and associated sensors in measuring key weather data. This includes data from three rotary-wing and two fixed-wing sUAS and included two commercially available systems and three university-developed research systems. Flight data were compared to regular radiosondes launched at the flight location, tower observations, and intercompared with data from other sUAS platforms. All platforms were shown to measure atmospheric state with reasonable accuracy, though there were some consistent biases detected for individual platforms. This information can be used to inform future studies using these platforms and is currently being used to provide estimated error covariances as required in support of assimilation of sUAS data into weather forecasting systems.

54 ENVIRONMENTAL SCIENCES↗

Synthesis of observed and simulated rain microphysics to inform a new Bayesian statistical framework for microphysical parameterization in climate models (Final Scientific Report)

This project aimed to investigate a new approach to bulk microphysics schemes. We identified the need to move beyond the fixed structural assumptions and approximation of existing schemes. For example, most schemes assume some functional form for the rain drop size distribution (e.g. a gamma or exponential distribution). Most schemes then evolve some number of statistical moments of that distribution via various processes, such as evaporation, sedimentation, collision-coalescence, and collisional breakup. These microphysical processes, in turn, are typically some fixed functional form derived from either some other (more detailed) model, or using some single-particle rates that are then integrated over the size distribution. While these process rate functions may have some free parameters to adjust, in other cases doing so is impossible (e.g. when the process rates are analytical or piecewise solutions to some target function). Our goal was to build and test a scheme that assumed no size distribution form, and used a series of power laws as the basis for the process rates. Moments of the size distribution would be predicted, but no underlying size distribution would be specified. Any number or choice of size distribution moments could be used, and any number of power laws could be employed to model the process rates. Thus, our approach is flexible, and can seamlessly scale across levels of complexity, as demanded by the data. Bayesian inference then would provide the formalism to estimate the model parameters and structure, allowing for robust uncertainty quantification. We proposed testing this framework in idealized simulations using bin microphysical schemes as a data source. We also proposed using real data to inform our microphysics scheme, and also that we would integrate our scheme into WRF.

58 GEOSCIENCES↗

Visualization for Scientific Discovery, Decision-Making, and Communication

Visualization–the use of visual elements to explore data, form hypotheses, or convey conclusions–is an integral part of the scientific process. Starting from an initial exploration of new data to illustrating outcomes to the general public, visualization is one of the most intuitive and powerful modes of communication. With the explosion of new data sources and types, unprecedented volumes of data, and new technologies, such as virtual reality and AI, visualization has become increasingly essential but also ever more challenging. Department of Energy’s (DOE) Office of Advanced Scientific Computing Research (ASCR) sponsored a Basic Research Needs workshop in January 2022 to understand the major opportunities and grand challenges in visualization tools and technologies for scientific computing, with a special focus on DOE-relevant applications and goals. The workshop identified five priority research directions (PRDs) for visualization to support scientific discovery, decision-making, and communication.

97 MATHEMATICS AND COMPUTING↗

Report for the ASCR Workshop on Visualization for Scientific Discovery, Decision-Making, and Communication

Visualization—the use of visual elements to explore data, form hypotheses, or convey conclusions—is an integral part of the scientific process. Starting from an initial exploration of new data to illustrating outcomes for the general public, visualization is one of the most intuitive and powerful modes of communication. With the explosion of new data sources and types, unprecedented volumes of data, and new technologies, such as virtual reality (VR) and artificial intelligence (AI), visualization has become increasingly essential but also ever more challenging. The Department of Energy’s (DOE) Office of Advanced Scientific Computing Research (ASCR) sponsored a Basic Research Needs workshop in January 2022 to understand the major opportunities and grand challenges in visualization tools and technologies for scientific computing as well as for DOE-relevant applications and goals in general. The workshop identified five priority research directions (PRDs) for visualization to support scientific discovery, decision making, and communication. The first three PRDs describe interconnected research themes addressing the need for new techniques to deal with complex data, uncertainty, and interpretability (PRD 1); the need for scalable and interoperable software stacks (PRD 2); and the challenges and opportunities inherent in new technologies, such as VR, cloud, or exascale computing (PRD 3). The remaining two PRDs describe foundational research themes that recognize the potential of visualizations to provide equitable access to information and to strengthen the scientific discourse (PRD 4); and the need to consider human factors when designing visualizations (PRD 5). Collectively, these PRDs form the pillars for a coherent, long-term research and development strategy in Visualization for Scientific Discovery, Decision-Making, and Communication in the context of the Office of Science’s mission scope.

97 MATHEMATICS AND COMPUTING↗

Requirements for Cataloging Hanford Geophysical Datasets

Environmental management activities at the Hanford Site produce extensive data about site conditions, contaminants, cleanup, and more. Managing and archiving that data requires a high degree of collaboration among site contractors and a high level of awareness by project managers and staff. Part of that effort is developing a Hanford Environmental Information and Data Index (HEIDI) to organize the data and maximize its value by making it findable and available for reuse. The objective is to catalog the disparate data sets collected to address the evolving needs of planning, executing, and documenting cleanup over several decades up to the present day, including links to active data sources when available. A properly implemented data catalog makes finding environmental datasets related to an area or theme a routine, reliable process, without requiring the searcher to have special knowledge that a data set exists and where it may be stored. In this project, a working group, including the U.S. Department of Energy, the Hanford Site contractors, and Pacific Northwest National Laboratory staff, identified needs and requirements for handling complex site data. Geophysical data was chosen as a test case because it can be large and complex and often involves multiple processing steps to extract the information incorporated into deliverables. The ability to document those steps was one of the requirements identified for the catalog. In addition to developing requirements, other activities included selecting a metadata schema and initial testing with the objective of determining whether the workflow and capabilities of selected data catalog software platforms were sufficient to implement and impose the identified requirements. This initial testing involved running the default catalog instance using the software platform of interest and altering the configuration to achieve each requirement, if possible. Where configuration alone was insufficient, the possibility of modifying the software by changing the code was examined, but not implemented. A follow-on task is planned to reprogram the code as necessary to implement requirements in a prototype catalog.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

National Energy Water Treatment & Speciation (NEWTS): A Water & Critical Mineral Database and Dashboard

The scarcity of water resources, the need for beneficial water reuse, and the challenges of wastewater treatment are becoming increasingly pressing in economic, social, and environmental domains. Addressing these concerns requires effective treatment strategies to manage wastewater streams and tackle environmental and economic issues. Furthermore, the recovery of critical minerals from the waste streams associated with energy production holds the promise of offsetting treatment costs and securing local sources of valuable minerals. However, relevant data on these waste streams are dispersed and challenging to locate. The process of ingesting such data into modeling software often involves multiple steps, requiring data restructuring to meet software-input requirements. The non-standardized reporting of water data makes data aggregation and reformatting a time-consuming process. Additionally, essential attributes necessary for modeling water treatment and mineral scale formation are frequently missing. Moreover, data gaps vary depending on the region of interest. Consequently, there is a pressing need for high-quality energy-water composition data that can be easily imported into water chemistry modeling software. To address this need, the National Energy Technology Laboratory has created the National Energy Water Treatment and Speciation (NEWTS) Database and Dashboard—a free online tool catering to community leaders and water researchers. NEWTS facilitates a comprehensive understanding of the composition of energy-related wastewater streams in the United States. The datasets provide detailed concentrations and speciation of major and minor aqueous compounds in energy-related wastewater streams, including power plant leachate, acid mine drainage, brackish water, and oil and gas produced water across the United States. Many of the aqueous species are critical minerals (Li, REEs) in high demand to modernize the world’s energy infrastructure. Many of the datasets also contain volumetric flow-rates needed to model the treatment and reuse scenarios in advanced aqueous chemistry software programs. The NEWTS Database and Dashboard offer public access to hitherto challenging-to-access datasets, presented in a standardized format that is tailored for easy input into aqueous chemistry modeling software. By performing the work needed to transform dispersed, disparate data sources into unified, model-ready datasets, NEWTS serves as an essential resource in advancing water treatment research and sustainable water resource management.

produced water management↗

Transfer Learning Trained LSTM Models for Household Load Profile Forecasting

Grid edge renewable energy resources, such as rooftop solar photovoltaics, closely interact with consumer load profiles. Therefore, forecasting future electricity demand, ideally at the individual household level, is indispensable. In this paper, we present a transfer learning enhanced household load profile forecasting method. First, we tune a long short-term memory forecasting model to perform day-ahead prediction of household electricity load profiles. Then we improve these individualized models using transfer learning, and we use k-means clustering to create optimal source data sets. We find average improvements of 4.38% (largest improvement of 10.71%) when the entire data set was used to train the source model and 2.45% (largest improvement of 11.57%) in the mean absolute error when households were first clustered and used to train separate source models for each cluster. We find that transfer learning with clustered data can effectively boost the forecasting performance of the LSTM models. We use realistic household power measurements for 148 real residential households in Austin, Texas.

deep learning↗

Comprehensive GOM Federal Waters Platform, Incident, Metocean, and Geohazard Dataset

The dataset contains integrated data from an array of disparate data sources, all spatially and temporally linked to platforms in the federal waters of the Gulf of Mexico (platform data from BSEE, 2020). Integrated data includes past reported incidents dating back to 1956 (BSEE, BOEM, MMS), metocean data (see Nelson et al. in review for source information), and geohazard data (see Nelson et al. in review for source information). Proprietary well production information was redacted from this dataset, but was used in resulting analytics.

Full System↗

SIDDA: SInkhorn Dynamic Domain Adaptation for image classification with equivariant neural networks

Modern neural networks (NNs) often do not generalize well in the presence of a ‘covariate shift’; that is, in situations where the training and test data distributions differ, but the conditional distribution of classification labels given the data remains unchanged. In such cases, NN generalization can be reduced to a problem of learning more robust, domain-invariant features. Domain adaptation (DA) methods include a broad range of techniques aimed at achieving this; however, these methods have struggled with the need for extensive hyperparameter tuning, which then incurs significant computational costs. In this work, we introduce SInkhorn Dynamic Domain Adaptation (SIDDA), an out-of-the-box DA training algorithm built upon the Sinkhorn divergence, that can achieve effective domain alignment with minimal hyperparameter tuning and computational overhead. We demonstrate the efficacy of our method on multiple simulated and real datasets of varying complexity, including simple shapes, handwritten digits, real astronomical observations, and remote sensing data. These datasets exhibit covariate shifts due to noise, blurring, differences between telescopes, and variations in imaging wavelengths. SIDDA is compatible with a variety of NN architectures, and it works particularly well in improving classification accuracy and model calibration when paired with symmetry-aware equivariant NNs (ENNs). We find that SIDDA consistently enhances the generalization capabilities of NNs, achieving up to a ${\approx}40\%$ improvement in classification accuracy on unlabeled target data, while also providing a more modest performance gain of $\lesssim 1\%$ on labeled source data. We also study the efficacy of DA on ENNs with respect to the varying group orders of the dihedral group DN, and find that the model performance improves as the degree of equivariance increases. Finally, if SIDDA achieves proper domain alignment, it also enhances model calibration on both source and target data, with the most significant gains in the unlabeled target domain—achieving over an order of magnitude improvement in the expected calibration error and Brier score. SIDDA’s versatility across various NN models and datasets, combined with its automated approach to domain alignment, has the potential to significantly advance multi-dataset studies by enabling the development of highly generalizable models.

79 ASTRONOMY AND ASTROPHYSICS↗

Data and Tools for Energy Planning and Analysis

The National Renewable Energy Laboratory (NREL) creates widely used data and tools to facilitate energy system planning and analysis. These software tools have been developed for complex research problems and perfected over real-world applications and laboratory validations. Some tools are award winners, others are open-source data explorers, and all are rigorously designed to empower decision-makers with accurate and accessible information. This software selection shows how NREL resources can help stakeholders achieve a clean, just, and resilient energy transformation.

data↗

Open data sets for assessing photovoltaic system reliability

Photovoltaic (PV) systems have become a cornerstone of renewable energy strategies, particularly due to the significant reduction in solar power costs over the past decade. However, the long-term reliability of PV installations presents a persistent challenge, requiring the development of advanced monitoring and predictive maintenance strategies. A wide range of data types is used to evaluate the health of PV systems, including environmental conditions, electrical performance, and inspection imagery. These data enable methodologies such as machine learning (ML) models for lifetime prediction and computer vision techniques for defect detection. However, the acquisition of high-quality and comprehensive data is difficult, particularly in terms of long-term consistency and data variety. Publicly available data sets serve as valuable resources for addressing these challenges, but they often suffer from fragmentation and are difficult to access. This paper presents a comprehensive review of existing open-source data sets related to PV degradation, analyzing their features, functionalities, and potential applications. We categorize these data sets based on the specific aspects of PV system information they cover, such as environmental conditions, operational monitoring, image inspection and module materials, and propose relevant tools and ML models for processing them. In addition, we propose practices for future data collection and usage, while also discussing potential directions in data-driven research. Our aim is to enhance data utilization and publication among researchers and industry professionals, promoting a deeper understanding of the role of data in enhancing the performance and durability of PV systems.

14 SOLAR ENERGY↗