Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Sources”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Fine-Root Ecology Database (FRED): A Global Collection of Root Trait Data with Coincident Site, Vegetation, Edaphic, and Climatic Data, Version 3.

To address the need for a centralized root trait database, we compiled the Fine-Root Ecology Database (FRED) from published and unpublished data sources. We have continued to add to the FRED database since the release of FRED 2.0 in 2018, and a new version of FRED is now available. FRED 3.0 has more than 150,000 observations of more than 330 root traits, with data collected from more than 1400 data sources. FRED 3.0 has 45% more root trait observations than FRED 2.0, particularly in the categories of root anatomy, morphology, and microbial associations; ancillary data on associated site, vegetation, edaphic, and climatic conditions from across the globe have also increased concurrently. FRED is focused on fine roots (traditionally defined as roots less than 2 mm in diameter), as coarse roots are studied using different methodology, often at very different scales, and have different traits and trait interpretations. However, FRED accepts data collected from roots of all sizes, and already contains several observations of coarse roots. Data collection will continue for the foreseeable future.

54 ENVIRONMENTAL SCIENCES↗

National Energy Water Treatment and Speciation (NEWTS) Database & Dashboard

The Department of Energy's Office of Fossil Energy & Carbon Management (DOE/FECM) through the National Energy Technology Laboratory (NETL) has launched a free online tool, the National Energy Water Treatment and Speciation (NEWTS) Database and Dashboard, which can be utilized by community leaders and water researchers to better understand the composition of energy-related wastewater streams. The NEWTS Database and Dashboard provide public access to difficult-to-access datasets, including the original data sources and the processed data forms for input into aqueous chemistry modeling software. The data provided by the tool will help mitigate environmental risks and identify possible sources of valuable critical minerals (CM). The goal of this ASME Power presentation is to highlight the data and capabilities of this free online-tool for obtaining high quality water datasets in formats that are easy for modeling the treatment and recovery of valuable resources from effluent waste stream associated with energy operations.

Siefert, Nicholas↗

Dynamic Line Rating Forecast Time Frames

As dynamic line ratings (DLR) are used to support different use cases, the preferred data source varies according to the forecast horizon. The below guidance is one possibility to utilize weather forecasting. Note that the time-periods of forecasts and spatial resolution may be subject to change as National Oceanic and Atmospheric Administration (NOAA)/National Weather Service (NWS) periodically upgrades its forecast models and data servers. Moreover, Persistence/ML Observations advances in analytics may justify combining data sources for more reliable DLR forecasts. The list below gives a time period (t) followed by possible guidance for data in that interval.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Decoding Ethiopian Abodes: Towards Classifying Buildings by Occupancy Type Using Footprint Morphology

Building occupancy classification plays a crucial role in urban planning, disaster management, and population modeling. Traditional methods often require extensive field surveys or detailed datasets, which can be time-consuming, expensive, and may yield incomplete or erroneous data. In this paper, we present a novel approach for classifying buildings as residential or non-residential using only building footprint data. By extracting geometric shape derivatives that characterize building morphology, we developed a high-accuracy classification model employing a combination of unsupervised and supervised learning methods. We utilized open-source data from Open Street Map, aggregating it to create binary labels for buildings based on their respective human use type. Our approach demonstrates the potential for scalability without the need for additional data sources other than building footprints and labels, offering a more efficient solution for building occupancy classification.

Adams, Daniel↗

A case study in contrastive learning information combination: Application to technical forensics of additive manufacturing filament source identification

Combination of information from disparate data sources into a single decision is a core challenge in many fields, including the field of technical forensics. Technical forensics (TF) utilizes technical characterization of questioned samples to determine properties of that sample; these properties are then used to infer information of forensic interest, such as provenance, age, or attribution. TF is utilized in traditional forensic applications, such as the attribution of material fragments from an explosive, and in nuclear forensic applications, such as the attribution of actinides which have been interdicted out of regulatory control. The challenge of combining information from disparate sources, described alternately by many terms including “Data Fusion” and “Data Integration”, is exacerbated in the technical forensics domain due to at least two factors: the challenge of interpreting each information source singularly, and the relatively small data set sizes available. Extensive literature exists attempting to combine technical forensics information sources, both in manual and automated processes. These attempts are often bespoke to the specific information sources (such as the bi-, tri-, or quad-isotope chart (Moody, Grant, and Hutcheon 2005)), with some emerging examples of simple early- and late- fusion (, respectively). Simultaneous to the information combination efforts described in the previous paragraph, the field of natural language processing attempted (and largely succeeded) in combining information from multiple non-technical information sources. The ecosystem of “multi-modal” language models, which can take text and images as input, and generate text and images as output, became large and diverse by 2025 (Khan et al. 2025). In a generalized sense, many of these methods are trained by learning neural networks which can convert raw text or images into a vector of numbers describing the text or image, hereafter called “embeddings” and the neural networks performing the conversion are called “embedders”. By using a separate embedder for text and images, finding coincident text and images (such as images with their captions), and optimizing the parameters of the embedders such that the embeddings for the text and the image are similar, the field has found a bridge between text and images (Girdhar et al. 2023). It is the contention of the authors of this report that this insight is not limited to text and images but instead can be extended to any modality which can be found coincidently. The subject of the rest of this report is the application of this method to example multi-modal technical forensic data. Some details about the data used in this report are not appropriate for this report, and are included in a companion report (PNNL-38669).

36 MATERIALS SCIENCE↗

A semantics-driven framework to enable demand flexibility control applications in real buildings

Decarbonising and digitalising the energy sector requires scalable and interoperable Demand Flexibility (DF) applications. Semantic models are promising technologies for achieving these goals, but existing studies focused on DF applications exhibit limitations. These include dependence on bespoke ontologies, lack of computational methods to generate semantic models, ineffective temporal data management and absence of platforms that use these models to easily develop, configure and deploy controls in real buildings. This paper introduces a semantics-driven framework to enable DF control applications in real buildings. The framework supports the generation of semantic models that adhere to Brick and SAREF while using metadata from Building Information Models (BIM) and Building Automation Systems (BAS). The work also introduces a web platform that leverages these models and an actor and microservices architecture to streamline the development, configuration and deployment of DF controls. The paper demonstrates the framework through a case study, illustrating its ability to integrate diverse data sources, execute DF actuation in a real building, and promote modularity for easy reuse, extension, and customisation of applications. The paper also discusses the alignment between Brick and SAREF, the value of leveraging BIM data sources, and the framework's benefits over existing approaches, demonstrating a 75% reduction in effort for developing, configuring, and deploying building controls.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Fine-Root Ecology Database (FRED): A Global Collection of Root Trait Data with Coincident Site, Vegetation, Edaphic, and Climatic Data, Version 4.

To address the need for a centralized root trait database, we compiled the Fine-Root Ecology Database (FRED) from published and unpublished data sources. We have continued to add to the FRED database since the release of FRED 1.0 in 2017, followed by 2.0 in 2018, and 3.0 in 2021. This new release of FRED 4.0 now has 213,941 observations of 238 root traits, for a combined total of roughly 3.4 million data fields for root traits and ancillary data together. FRED 4.0 has 39.8% more root trait observations than FRED 3.0 and a 34.4% increase in unique data sources. This release of FRED 4.0 also includes significant increases in geographic regions that have long been underrepresented in global datasets, notably in the tropical low latitudes. Ancillary data on associated site, vegetation, edaphic, and climatic conditions from across the globe have also increased concurrently with root trait observations. FRED is focused on fine roots (traditionally defined as roots less than 2 mm in diameter), as coarse roots are studied using different methodology, often at very different scales, and have different traits and trait interpretations. Despite this fine-root focus, FRED accepts data collected from roots of all sizes and contains observations of many root classes including coarse roots. Data collection will continue for the foreseeable future. The FRED4_Entire_Database_2026.csv file is the flat csv data file for FRED 4.0, and the FRED4_dd.csv file is the data dictionary of all columns available in FRED, including column IDs, column names, definitions, and unit (where applicable).

54 ENVIRONMENTAL SCIENCES↗

Natural Language Processing for Text Based Event Extraction: Identifying Events of Interest Related to Worldwide State-Sponsored Civil Nuclear Power

Beginning in FY20, SRNL was funded by the National Nuclear Security Administration’s Office of Defense Nuclear Non-Proliferation Research and Development to develop a prototype natural language processing/natural language understating machine learning-based modeling and analysis pipeline to extract and forecast events of interest from massive open data sources. The working hypothesis within the approach is that contextual shifts in key words and phrases act as indicators of events of interest over time. Therefore, by identifying points in time where contextual shifts occur, events of interest can be extracted along with explicit and implicit connections of entities and activities. The development of the preliminary prototype pipeline proved successful, meriting further testing of the pipeline on more broad topical domains and in a worldwide data environment. Therefore, SRNL, in collaboration with the Sanghani Center for Artificial Intelligence and Data Analytics at Virginia Tech, have continued development with a test case of identifying events of interest related to worldwide state-sponsored civil nuclear power in open data sources. In the first year of this follow-on effort, the team has curated domain-specific data corpuses using an automated scheme and applied the modeling and analysis pipeline. This robust, focused, and efficient approach consists of an ensemble of analyses applied to time dependent word embedding models that are trained on the data corpuses. In this report, the team has demonstrated the capability of the existing pipeline (as development has continued in parallel) by exploring several specific case-studies centered around Rosatom’s international activities regarding the planning, construction, operation, and/or shutdown of nuclear reactors. A basic timeline events has been generated by manually cataloging known “milestone” events that have occurred at reactors in Turkey, Finland, Hungary, and Egypt and compared with the output of the modeling pipeline. In this approach, the team has characterized the lead time using the prototype pipeline, as well as the ability to capture relevant information, which proved 100% successful. A deep dive example of the Akkuyu reactor (Turkey) is presented that shows the breadth of information that can be captured using the approach. In this case study, events were extracted pertaining to the planning/construction of Akkuyu including protests from the population, information campaigns in response to the protests, forged regulatory documents and lawsuits, budgetary/shareholder information, geopolitical tensions, and the various construction milestones. This has demonstrated the pipeline’s utility as a research aid or real-time event extraction tool, where summary-level information and detailed text extractions from millions of articles or Tweets across long time periods can be generated with significantly less effort than current techniques.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Open Source Evaluation Framework for Solar Forecasting

The Solar Forecast Arbiter is an open-source evaluation framework for solar forecasting. The framework enables evaluations of solar irradiance, solar power, and net-load forecasts that are impartial, repeatable and auditable. The Solar Forecast Arbiter addresses stakeholder-informed use cases including evaluation of forecast skill, comparisons to reference data sets, private forecast trials, and evaluation of probabilistic forecast skill. The framework includes a data validation toolkit, reference data sources, data privacy protocols, and benchmark forecast capabilities for intra-hour and day ahead forecast horizons. Reports and metrics communicate the relative merits of the test and benchmark forecasts. The reports are created from standardized templates and include graphics for qualitatively evaluating deterministic and probabilistic forecasts and standard metrics for quantitatively evaluating forecasts. The Solar Forecast Arbiter is designed to support all solar forecasting stakeholders, including Solar Forecasting 2 Topic Area 2 and Topic Area 3 teams.

14 SOLAR ENERGY↗

Fixed Target Serial Data Collection at Diamond Light Source

Serial data collection is a relatively new technique for synchrotron users. A user manual for fixed target data collection at I24, Diamond Light Source is presented with detailed step-by-step instructions, figures, and videos for smooth data collection.

Horrell, Sam↗

Lowering the barrier to access information-rich transient kinetic data for machine learning methods

Transient kinetic data contain a wealth of information about intrinsic features of a catalyst as well as the reaction mechanism. Currently, high volume transient data is underutilized, and data science methods could both increase the value of information that can be extracted from this data, integrate experimental with theoretical data sources, and accelerate the pace of catalyst technology advancement. Transient kinetic characterizations with simple probe molecules exhibiting reversible adsorption, irreversible adsorption and bulk-surface diffusion are presented as training components for similar experiments with more complex surface reactions. In conclusion, by increasing the availability and accessibility of transient kinetic data through details of its structure and acquisition, we aim to decrease the barrier for data scientists to apply machine learning methods to this valuable data source.

Catalysis↗

Understanding Event Trajectories Across Massive Temporal Datasets with Word Embeddings and Visualization

In collaboration with researchers from Virginia Tech, Savannah River National Laboratory has continued development of a natural language processing pipeline to identify and extract events of interest from massive open data sources in the domain of worldwide state-sponsored civil nuclear energy. The foundation of the pipeline is built on compass aligned temporal word embedding models, whereby contextual shifts are automatically identified by comparing keyword embedding vectors across successive time windows. Within the approach, a contextual shift indicates the occurrence of a potential event of interest. However, in such a broad topical domain that captures events at a global scale, across various life cycle stages, and across numerous different technology types, a user that is monitoring events may have broad interests in capturing many different event types with varying degrees of signal. As such, the quantity of information that may be returned from an automated event extraction pipeline can be substantial, requiring manual effort to sift through the information to identify any relevant bits of information. Therefore, a more streamlined workflow that aids in directing a user toward specific information at different points in time is necessary. The workflow presented here has been developed with this concept in mind, built on top of the initial prototype event extraction pipeline, whereby a user can analyze temporal text-based data sources at multiple different contextual levels to isolate key points in time and key subdomains captured within a data corpus. Using multiple corpuses that consist of approximately 7 million Tweets and 7 million news articles, the team has extended compass aligned temporal word embedding models to establish an interconnected and hierarchical structure that relates known key words of interest to documents, local topics (i.e., within a time window), and global topics across the corpuses. All of this information is packaged into a visual analytics system that is linked to the information extraction pipeline and enables a user to identify contextual information that describes the evolution of a high dimensional embedding space across time to isolate changes of interest and explore associated events. This report demonstrates the use of these analytics and a means to fuse information across multiple datasets.

97 MATHEMATICS AND COMPUTING↗

Sources of Propane Consumed in California

Project Scope: The objective of this study is to specify the sources of propane consumed in California. It answers the questions, where does the propane used in California come from and how was it produced? The results of this study provide comprehensive, transparent, and verifiable estimates, based on the 2018 market. The information provided in this report is suitable for use to assess the life cycle carbon intensity of propane used as a transportation fuel in California. As the 2009 Low Carbon Fuel Standard (LCFS) aims to reduce California’s greenhouse gas (GHG) emissions and other smog-forming and toxic air pollutants, the appropriate designation of carbon intensity for propane as a transportation fuel is important for evaluating propane’s potential to contribute to GHG goals and understandings in the context of various actions. This study focuses on estimating the shares of total propane consumed in the state of California produced from petroleum refineries, natural gas plants, and bituminous sands sources inside California and elsewhere. Results: An estimated 590 million gallons of propane were consumed in California in 2018, of which, 59.5% originated from refinery production and 40.5% originated from natural gas plants. The majority of this was sourced from refinery production in California, 334 million gallons. Most of the propane imported to California for consumption was sourced from natural gas plants, 113 million gallons, with over half of the imported volume sourced from Canada. The volume sourced from bituminous sand upgrader and fractionator operations was negligible. Details from this analysis are presented in the table below which provides an overview of the propane flows estimated in this study by region and production method. The shares and volumes presented here represent a snapshot for 2018. A significant increase in propane demand, such as could be caused by increased use of propane as a transportation fuel in the state, would affect California’s propane production, imports, and exports. The method and data sources used for the estimates provided in this report also provide the framework which could be used for future updates. Key Method Considerations: The values presented here are based on a two-step approach where the first step was to determine the flows of propane into and out of California from different regions and the second step was to estimate the propane production methods in each region. A volume balance approach is used as the primary method for tracking the volume of propane in and out of California as propane production and import volumes are available by Petroleum Administration of Defense District (PADD) from EIA and neither inter-PADD propane transfers nor state-specific non-prime supplier consumption are available from a public data source. The volume balance performed for this study covered PADD 5 (the West Coast), which includes Arizona, California, Nevada, Oregon, and Washington. The volume balance used all available public datasets to determine propane production, imports, exports, and consumption. Volumes unaccounted for by these datasets were estimated using the resulting volume balance by assuming market equilibrium. Consumption within each state in PADD 5 was estimated based on known import, export, and production volumes and this amount was used to develop the volume balance. EIA only tracks consumption at the state level by prime supplier sales. The volume balance approach provides the basis to correct for additional propane consumed in-states where propane is transferred to California. To determine the California propane sources and trade in 2018. volume of propane consumed in California, the volume balance approach is again used where it was estimated all imported volumes not specifically flagged for re-export were consumed, and the remaining consumption was produced in-state. The California Energy Commission (CEC) provided the total volume of propane imported and exported from California in 2018; this volume data set along with commodity tracking from the Canada Energy Regulator (CER) and the International Trade Commission (ITC) which tracks port of entry and final destination was used to determine where propane originated from and where it was ultimately consumed. For example, the CER tracks propane leaving Canada and entering each state within the U.S. Imported propane from Canada to California – marked for California – is assumed to be consumed in California. When no further data were available, import volumes were assumed to be consumed in California without pass-through (i.e., no propane imported to California was directly sold and exported). In most cases, the production method for each propane source region was applied to the volume of propane transferred to California. In other words, the shares of propane sourced from natural gas and refineries for each production region was assigned to California imports based on their contribution to the total volume flows into California to determine the production method for propane consumed in-state. For volumes imported into California from PADD 4, Washington State, Canada, and the rest of the world (Argentina, Chile, Norway, Peru, South Korea, and Trinidad and Tobago), the volumes sourced from petroleum refineries and natural gas plants reflect the either production ratio for the region or, in cases where the sources specific to the amounts exported to California could be determined, the sources specific to the volumes transferred to California.

03 NATURAL GAS↗

Open Data for Nuclear Explosion Monitoring (NEM) [Slides]

The data sources tend to have the highest quality data and metadata, particularly for more recent data sets. Early data from sources such as IRIS tend to have some metadata issues.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

NetGraf: An End-to-End Learning NetworkMonitoring Service (NetGraf) v1

NetGraf is a novel end-to-end learning monitoring system that utilizes current monitoring tools, merges multiple data sources into one dashboard for easy use, and provides machine learning libraries to analyze the data and perform real-time anomaly findings. Using a database backend, NetGraf can learn performance trends and show users if network performance has degraded. We demonstrate how NetGraf can easily be deployed through automation services and linked to multiple monitoring sources to collect data. Via the machine learning innovation and merging various data sources, NetGraf aims to fulfill the need for holistic learning network telemetry monitoring. To the best of our knowledge, this is the first-ever end-to-end learning monitoring service. We demonstrate its use on two network setups to showcase its impact.

Mohammed, Bashir↗

Data and Scripts Associated with "Modeling Ecohydrological Responses of Vegetation to Urban Microclimates Using the E3SM Land Model"

This dataset supports the study of vegetation ecohydrological responses to urban microclimates using the land component of the Energy Exascale Earth System Model (ELM) at four urban sites in Knoxville, Tennessee, USA. It includes the model inputs, simulation outputs, and associated scripts for running ELM simulations and analyzing the resulting data. The Model_Inputs folder includes static surface data, satellite-derived phenology (i.e., leaf area index), and atmospheric forcing data used to drive ELM simulations. Detailed descriptions of these datasets are provided in Section 2.3.2 of the associated manuscript. The Model_Outputs folder contains simulation results for the baseline, treatment, and ensemble experiments. Outputs from the baseline and treatment simulations are provided as raw ELM NetCDF files. Because the raw outputs from the 4,000-member ensemble are prohibitively large, the ensemble results are provided as summarized CSV files, which also serve as the source data for Figure 5 of the associated manuscript. The Scripts folder contains three components: E3SM, the core codebase of the Energy Exascale Earth System Model (E3SM); elm-olmt, the Offline Land Model Testbed (OLMT) used to perform the simulations; and knoxville_elm, which contains the analysis scripts used to process model outputs and generate the figures and results presented in the associated manuscript. Additional information is provided in Scripts_readme.txt within the Scripts directory.

Lu, Xiaoman [ORNL] (ORCID:0000000306698780)↗

Adjoint-Based Inversion of Geodetic Data for Sources of Deformation and Strain

An adjoint-based formulation leads to a particularly efficient approach for inverting geodetic measurements for the source of the deformation. Specifically, the quantities necessary to iteratively improve the fit to the observations can be computed with just three forward calculations, one to obtain the current residuals, another to solve the adjoint problem, and a third to compute the step length. An inversion algorithm utilizing the adjoint-based gradient is applied to a set of Interferometric Synthetic Aperture Radar (InSAR) data gathered between 2016 and 2018 over the Tulare Basin in California's Central Valley. Because the measured deformation is due to groundwater withdrawal, a penalty function is included in the inversion to avoid placing aquifer volume change in locations that are far from any documented wells. The solution of the inverse problem provides estimates of aquifer compaction that provide a match to the observed range changes while honoring the well data. The solution indicates an average aquifer volume loss of 2.17 km 3 /year over the two year period from January 2016 to January 2018, encompassing one drought year (2016) and one wet year (2017). Finally, this magnitude of lost volume is compatible with the 3.1 km 3 /year decrease in water volume for the entire Central Valley, estimated from GRACE satellite gravity data.

58 GEOSCIENCES↗