Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Sources”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

BESS-STAIR: a framework to estimate daily, 30m, and all-weather crop evapotranspiration using multi-source satellite data for the US Corn Belt

Abstract. With increasing crop water demands and drought threats, mapping andmonitoring of cropland evapotranspiration (ET) at high spatial and temporalresolutions become increasingly critical for water management andsustainability. However, estimating ET from satellites for precise waterresource management is still challenging due to the limitations in bothexisting ET models and satellite input data. Specifically, the process of ETis complex and difficult to model, and existing satellite remote-sensing datacould not fulfill high resolutions in both space and time. To address theabove two issues, this study presents a new high spatiotemporal resolution ETmapping framework, i.e., BESS-STAIR, which integrates a satellite-drivenwater–carbon–energy coupled biophysical model, BESS (Breathing Earth SystemSimulator), with a generic and fully automated fusion algorithm, STAIR(SaTallite dAta IntegRation). In this framework, STAIR provides daily 30'mmultispectral surface reflectance by fusing Landsat and MODIS satellite datato derive a fine-resolution leaf area index and visible/near-infrared albedo,all of which, along with coarse-resolution meteorological and CO 2 data, are used to drive BESS to estimate gap-free 30 m resolution daily ET.We applied BESS-STAIR from 2000 through 2017 in six areas across the US CornBelt and validated BESS-STAIR ET estimations using flux-tower measurementsover 12 sites (85 site years). Results showed that BESS-STAIR daily ETachieved an overall R2=0.75, with root mean square error RMSE=0.93 mm d -1 and relative error RE =27.9 % when benchmarkedwith the flux measurements. In addition, BESS-STAIR ET estimations capturedthe spatial patterns, seasonal cycles, and interannual dynamics well indifferent sub-regions. The high performance of the BESS-STAIR frameworkprimarily resulted from (1) the implementation of coupled constraints onwater, carbon, and energy in BESS, (2) high-quality daily 30 m data from theSTAIR fusion algorithm, and (3) BESS's applicability under all-skyconditions. BESS-STAIR is calibration-free and has great potentials to be areliable tool for water resource management and precision agricultureapplications for the US Corn Belt and even worldwide given the globalcoverage of its input data.

54 ENVIRONMENTAL SCIENCES↗

Sharing the Sun: Community Solar Deployment and Subscriptions (As of January 2026)

The community solar market analysis presented here is based primarily on data collected through Sharing the Sun, an initiative of the National Community Solar Partnership+ (NCSP+). Sharing the Sun data collection and analysis are conducted by the National Laboratory of the Rockies (NLR) as part of its support for implementation of NCSP+. NLR first released a dataset of community solar projects in 2018 and updates it biannually. The January 2026 dataset, data collection methodology, and all the previous datasets are available from NLR's Data Catalog: https://data.nlr.gov/submissions/244. The dataset presents project-level information including location, capacity, operating utility, and year of interconnection. The dataset is created from multiple data sources such as utility data, public utility commissions, project developer websites, media releases, primary data collection by NLR, and data provided by developers under nondisclosure agreements. This presentation builds on a previous analysis of the community solar project dataset, Sharing the Sun: Community Solar Deployment and Subscriptions (as of June 2024). Dr. Gabriel Chan and his team at the University of Minnesota contribute to this effort. NCSP+ is led and funded by U.S. Department of Energy's Integrated Energy Systems Office (IESO).

14 SOLAR ENERGY↗

COVID-19, An Exercise in Data Governance at Sandia National Laboratories

In April of 2020, Sandia National Laboratories had an urgent need to identify and manage the data that could be used to create mobile applications, models, reports, and visualizations to assist management in safely bringing the workforce onsite during the COVID-19 pandemic. Multiple divisions volunteered to design and build software solutions; meanwhile, requests for new data sources, including duplicate requests, were inundating Information Technology (IT) and data owners. The Enterprise Data Governance Team was assigned to resolve obtaining and accessing new sources of data in an accelerated timeframe. Through successful collaboration with multiple stakeholders and domain owners across Sandia, the Enterprise Data Governance Team rapidly developed a centralized data strategy and solution for use in safeguarding the Sandia workforce during the COVID-19 pandemic. This foundation enabled teams to successfully develop solutions, including reports for executives and management as well as the data for modeling and scientific analysis.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Chimera D-Series Gravitational Wave Emission Sourced from Neutrino Anisotropy

Gravitational wave data sourced from the time-dependent anisotropic neutrino emission, as well as the time-dependent fluid quadrupole motion, in the Chimera D-Series three-dimensional core collapse supernova simulations. Data from three models initiated from three different progenitors are presented: D9.6-3D, D15-3D, and D25-3D. Please see the README for more information about the data structure and progenitors.

79 ASTRONOMY AND ASTROPHYSICS↗

Leveraging Large Language Models for Real-World Data Evidence: A Framework for Automated Treatment Extraction and Data Harmonization

Background: The ability to comprehensively collect treatment information from cancer patient medical records would enable studies to evaluate real-world benefits and risks tied to specific treatments. Currently, it is difficult to system- atically collect high-quality treatment information because it is often stored in unstructured text. Manually extracting and standardizing drug and regimen data is time-intensive. Recent advances in large language models (LLMs) offer a potential solution for automated extraction of structured treatment information from clinical text. Objective: This study systematically evaluates the utility of four LLMs from the Llama family for automated extraction of oncology treatment information from clinical text. This information can guide researchers using cancer registry data to provide insights into cancer care and outcomes beyond clinical trials. Methods: Four instruction-tuned Llama models with varying parameter counts (1B, 3B, 8B, and 70B) were evaluated for their ability to extract treatment information from clinical documents. A unified oncology knowledge base integrating seven major public data sources was developed to standardize and normalize extracted entities—a critical step for harmonizing data from diverse sources. Extracted treatment data were compared against expert-annotated ground truth. Model performance was assessed using accuracy metrics (Precision, Recall, F1-Score) and opera- tional feasibility metrics, including processing speed and structural compliance of the output. Results: A strong positive correlation was observed between model size and extraction accuracy. F1-score improved from 0.609 for the 1B model to 0.710 (3B), 0.807 (8B), and 0.828 (70B). While larger models demonstrated superior accuracy and compliance, they incurred higher computational costs. The modest performance difference between 8B and 70B suggests diminishing returns with increasing model size. Conclusions: LLMs represent a viable technology for automating oncology treatment extraction. The 8B-parameter model emerged as a highly effective option, balancing high accuracy and computational efficiency. Selecting an appropriate LLM for deployment in cancer registries involves a trade-off between desired accuracy and available operational resources. Harmonizing extracted entities with the oncology knowledge base facilitates standardized integration into common data models, enhancing data quality for real-world evidence analyses.

artificial intelligence↗

Event Definition for the Automated Detection of Nuclear Proliferation Activities

In FY2020, Savannah River National Laboratory (SRNL) in collaboration with the Discovery Analytics Center (DAC) at Virginia Polytechnic Institute and State University (VT) began developing a demonstration prototype system that uses multiple machine learning and data analytic methods on largescale open data sources to identify new, developing, or undeclared nuclear programs. One of the most challenging aspects of applying machine learning techniques to such a problem is the high likelihood of extremely sparse data from disparate sources. To overcome this challenge, the current work will use a strategic combination of supervised, semi-supervised, and unsupervised learning techniques to ingest and fuse data streams to make a forecast of nuclear activities in a targeted geospatial location. Identifying potential data sources and training supervised learning algorithms is dependent upon the development of a robust foundation of targeted event domains that fundamentally define the nuclear activities of interest. This report documents the definition of a hierarchical structure for both nuclear activity and event domains that will be used to guide the research team in development or use of existing semantic dictionaries that are instrumental to searching, parsing, and categorizing events for the forecasting system’s use.

97 MATHEMATICS AND COMPUTING↗

Building Performance Database API (BPD API) v2.1

The Building Performance Database (BPD) is the largest publicly-available source of measured energy performance data for buildings in the United States. It contains information about the building's energy use, location, and physical and operational characteristics. The BPD can be used by building owners, operators, architects and engineers to compare a building's energy efficiency against customized peer groups, identify energy efficiency opportunities, and set energy efficiency targets. It can also be used by energy efficiency program implementers and policymakers to analyze energy efficiency features and trends in the building stock. The BPD compiles data from various data sources, converts it into a standard format, cleanses and quality checks the data, and provides users with access to the data in a way that maintains anonymity for data providers. This software is the database and the Application Programming Interface (API). Users can utilize the BPD's data to develop their own applications using the API. Version 2.1 included a major update for multiple years of data and refactoring of code for faster queries.

Mathew, Paul↗

C-HER Metadata Overview: Approach, Standards, and Rigor for the Centralized Health and Exposomic Resource

The Centralized Health and Exposomic Resource (C-HER) unifies environmental, demographic, geographic, and health-related data for exposomic research. The source data differ in format, geographic coverage, time period, resolution, terminology, and documentation. We use a common metadata framework to describe those differences and to record how each data resource has been processed, documented, and ingested. This document relates only to the C-HER metadata framework. It explains the information that is recorded for each resource, the standards used to organize that information, the conditions for metadata completeness, and the relationship between metadata and quality review. It is intended for those who need to understand what C-HER metadata communicates and how it supports appropriate use of the data. It is not an implementation specification or procedure. It does not document the database schema, source code, deployment configuration, transformation algorithms, or dataset-specific QA/QC thresholds. Those materials are maintained separately.

MacFarland, Midgie [ORNL] (ORCID:0009000807354078)↗

Exploring Advanced Computational Tools and Techniques with Artificial Intelligence and Machine Learning in Operating Nuclear Plants

This report presents the project Idaho National Laboratory conducted for Nuclear Regulatory Commission to explore the advanced computational tools and techniques, such as artificial intelligence (AI) and machine learning (ML), for operating nuclear plants. The report reviews the nuclear data sources, with the focus on the operating experience data, that could be applied by advanced computational tools and techniques. Plant-specific and generic (national and international) data from different sources are described. The report describes the relationships between statistics and AI/ML and then introduces the most widely used AI/ML algorithms in both supervised and unsupervised learning. The report reviews the recent applications of advanced computational tools and techniques in various fields of nuclear industry, such as reactor system design and analysis, plant operation and maintenance, and nuclear safety and risk analysis. Finally, the report presents the insights from the project on the potential applicability of AI/ML techniques in improving advanced computational capabilities, how the advanced tools and techniques could contribute to the understanding of safety and risk, and what information would be needed to provide meaningful insights to decision makers. The report also documents an NRC survey on the current state of commercial nuclear power operations relative to the use of AI and ML tools as well as the role of AI/ML tools in nuclear power operations was published by the NRC as in FRN NRC-2021-0048 in April 2021. A summary of the survey including the survey questions, survey participants, survey responses, and the conclusions and insights derived from the survey is provided in the report. Finally, the report investigates potential applications of using AI/ML in operating NPPs and advanced reactors (both advanced LWRs and advanced NLWRs) to improve nuclear plant safety and efficiency. Three main application fields are defined and discussed: (1) plant safety and security assessments; (2) plant degradation modeling, fault and accident diagnosis and prognosis; and (3) plant operation and maintenance efficiency improvement.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Open Science principles for accelerating trait-based science across the Tree of Life

Synthesizing trait observations and knowledge across the Tree of Life remains a grand challenge for biodiversity science. Species traits are widely used in ecological and evolutionary science, and new data and methods have proliferated rapidly. Yet accessing and integrating disparate data sources remains a considerable challenge, slowing progress toward a global synthesis to integrate trait data across organisms. Trait science needs a vision for achieving global integration across all organisms. In this perspective, we outline how the adoption of key Open Science principles—open data, open source and open methods—is transforming trait science, increasing transparency, democratizing access and accelerating global synthesis. To enhance widespread adoption of these principles, we introduce the Open Traits Network (OTN), a global, decentralized community welcoming all researchers and institutions pursuing the collaborative goal of standardizing and integrating trait data across organisms. We demonstrate how adherence to Open Science principles is key to the OTN community and outline five activities that can accelerate the synthesis of trait data across the Tree of Life, thereby facilitating rapid advances to address scientific inquiries and environmental issues. Lessons learned along the path to a global synthesis of trait data will provide a framework for addressing similarly complex data science and informatics challenges.

59 BASIC BIOLOGICAL SCIENCES↗

Carbon Storage Open Database

The Carbon Storage Open Database is a collection of spatial data obtained from publicly available sources published by several NATCARB Partnerships and other organizations. The carbon storage open database was collected from open-source data on ArcREST servers and websites in 2018, 2019, 2021, and 2022. The original database was published on the former GeoCube, which is now EDX Spatial, in July 2020, and has since been updated with additional data resources from the Energy Data eXchange (EDX) and external public data resources. The shapefile geodatabase is available in total, and has also been split up into multiple databases based on the maps produced for EDX spatial. These are topical map categories that describe the type of data, and sometimes the region for which the data relates. The data is separated in case there is only a specific area or data type that is of interest for download. In addition to the geodatabases, this submission contains: 1. A ReadMe file describing the processing steps completed to collect and curate the data. 2. A data catalog of all feature layers within the database. Additional published resources are available that describe the work done to produce the geodatabase: Morkner, P., Bauer, J., Creason, C., Sabbatino, M., Wingo, P., Greenburg, R., Walker, S., Yeates, D., Rose, K. 2022. Distilling Data to Drive Carbon Storage Insights. Computers & Geosciences. https://doi.org/10.1016/j.cageo.2021.104945 Morkner, P., Bauer, J., Shay, J., Sabbatino, M., and Rose, K. An Updated Carbon Storage Open Database - Geospatial Data Aggregation to Support Scaling -Up Carbon Capture and Storage. United States: N. p., 2022. Web. https://www.osti.gov/biblio/1890730 Morkner, P., Rose, K., Bauer, J., Rowan, C., Barkhurst, A., Baker, D.V., Sabbatino, M., Bean, A., Creason, C.G., Wingo, P., and Greenburg, R. Tools for Data Collection, Curation, and Discovery to Support Carbon Sequestration Insights. United States: N. p., 2020. Web. https://www.osti.gov/biblio/1777195 Disclaimer: This project was funded by the United States Department of Energy, National Energy Technology Laboratory, in part, through a site support contract. Neither the United States Government nor any agency thereof, nor any of their employees, nor the support contractor, nor any of their employees, makes any warranty, express or implied, or assumes any legal liability or responsibility for the accuracy, completeness, or usefulness of any information, apparatus, product, or process disclosed, or represents that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by trade name, trademark, manufacturer, or otherwise does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States Government or any agency thereof. The views and opinions of authors expressed herein do not necessarily state or reflect those of the United States Government or any agency thereof.

carbon storage↗

National and subnational short-term forecasting of COVID-19 in Germany and Poland during early 2021

During the COVID-19 pandemic there has been a strong interest in forecasts of the short-term development of epidemiological indicators to inform decision makers. In this study we evaluate probabilistic real-time predictions of confirmed cases and deaths from COVID-19 in Germany and Poland for the period from January through April 2021. We evaluate probabilistic real-time predictions of confirmed cases and deaths from COVID-19 in Germany and Poland. These were issued by 15 different forecasting models, run by independent research teams. Moreover, we study the performance of combined ensemble forecasts. Evaluation of probabilistic forecasts is based on proper scoring rules, along with interval coverage proportions to assess calibration. The presented work is part of a pre-registered evaluation study. We find that many, though not all, models outperform a simple baseline model up to four weeks ahead for the considered targets. Ensemble methods show very good relative performance. The addressed time period is characterized by rather stable non-pharmaceutical interventions in both countries, making short-term predictions more straightforward than in previous periods. However, major trend changes in reported cases, like the rebound in cases due to the rise of the B.1.1.7 (Alpha) variant in March 2021, prove challenging to predict. Multi-model approaches can help to improve the performance of epidemiological forecasts. However, while death numbers can be predicted with some success based on current case and hospitalization data, predictability of case numbers remains low beyond quite short time horizons. Additional data sources including sequencing and mobility data, which were not extensively used in the present study, may help to improve performance.

60 APPLIED LIFE SCIENCES↗

UMLS users and uses: a current overview

Abstract The US National Library of Medicine regularly collects summary data on direct use of Unified Medical Language System (UMLS) resources. The summary data sources include UMLS user registration data, required annual reports submitted by registered users, and statistics on downloads and application programming interface calls. In 2019, the National Library of Medicine analyzed the summary data on 2018 UMLS use. The library also conducted a scoping review of the literature to provide additional intelligence about the research uses of UMLS as input to a planned 2020 review of UMLS production methods and priorities. 5043 direct users of UMLS data and tools downloaded 4402 copies of the UMLS resources and issued 66 130 951 UMLS application programming interface requests in 2018. The annual reports and the scoping review results agree that the primary UMLS uses are to process and interpret text and facilitate mapping or linking between terminologies. These uses align with the original stated purpose of the UMLS.

Amos, Liz↗

GeoAI for Public Health

Infectious disease spread within the human population can be conceptualized as a complex system composed of individuals who interact and transmit viruses through spatio-temporal processes that manifest across and between scales. The complexity of this system ultimately means that the spread of infectious diseases is difficult to understand, predict, and respond to effectively. Research interest in GeoAI for public health has been fueled by the increased availability of rich data sources such as human mobility data, OpenStreetMap data, contact tracing data, symptomatic online surveys, retail and commerce data, genomics data, and more. This data availability has resulted in a wide variety of data-driven solutions for infectious disease spread prediction which show potential in enhancing our forecasting capabilities. This chapter (1) motivates the need for AI-based solutions in public health by showing the heterogeneity of human behavior related to health, (2) provides a brief survey of current state-of-the-art solutions using AI for infectious disease spread prediction, (3) describes a use-case of using large-scale human mobility data to inform AI models for the prediction of infectious disease spread in a city, and (4) provides future research directions and ideas.

Zufle, Andreas↗

Development of National New Construction Weighting Factors for the Commercial Building Prototype Analyses (2003-2018)

The U.S. Department of Energy (DOE), Office of Building Technologies tasked Pacific Northwest National Laboratory to develop construction weights for various commercial building categories for the purpose of estimating weighted national energy savings from the development of ANSI/ASHRAE/IESNA Standard 90.1-2019 compared to ANSI/ASHRAE/IESNA Standard 90.1-2016. Disaggregate construction volume data was acquired from the McGraw Hill Construction Database for the years 2003-2018 and analyzed to develop detailed construction weights by climate zones, subzones and by states. These weights are provided in this report and will be subsequently used in developing a weighted national energy savings estimate for the impact of the 2019 standard. For the current update, PNNL reviewed the same data source with the latest construction data for the years 2003-2018.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Review of the Nuclear Energy Agency (NEA) Ancillary Thermodynamic Database (TDB) Volume (DRAFT REV. 0)

The Nuclear Energy Agency (NEA) Ancillary data volume comprises thermodynamic data of mineral and aqueous species that, in addition to Auxiliary Data (as referred to in previous NEA thermodynamic data volumes), is necessary to calculations of chemical interactions relevant to radioactive waste management and nuclear energy. This SAND report is a review of the NEA Ancillary data critical reviews volume of thermodynamic data parameters. The review given in this report mainly involves data comparison with other thermodynamic data assessments, analysis of thermodynamic parameters, and examination of data sources. Only new and updated data parameters were considered in this review. Overall, no major inconsistencies or errors were found as allowed by the comparisons conducted in this review. Some remarks were noted, for example, on the consideration of relevant studies and/or comparisons on the analysis and retrieval of thermodynamic data parameters not cited in the respective sections.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Development of National New Construction Weighting Factors for the Commercial Building Prototype Analyses (2008-2022)

The U.S. Department of Energy (DOE) tasked Pacific Northwest National Laboratory (PNNL) with updating commercial building construction weights for the purpose of estimating national and state-by-state energy savings impacts of changes made to various commercial energy codes and standards. A similar activity was last completed by PNNL in 2020 using disaggregate construction volume data acquired from the Dodge Data & Analytics database (formerly McGraw Hill) for the years 2003-2018 (Lei et al, 2020). As time passes, changes in economic and social demand reshape construction volume trends. For the current update, PNNL reviewed the same data source with the latest construction data for the years 2008-2022. For commercial building analyses, PNNL typically uses a suite of 16 prototype buildings simulated in the 19 ASHRAE climate zones with 16 of them present in the United States. The 2008-2022 commercial building weighting factors were derived using the same approach employed to develop the 2003-2018 set (Lei et al, 2020). Applying the construction volume data from the database to the prototypes and climate zones resulted in the following new construction area-based weighting factors. Table ES.1 shows the weighting factors including all building categories found in the database, and Table ES.2 shows the weighting factors normalized to include only buildings represented by the 16 prototypes. Section 3.0 also includes national- and state-level weighting factors by area and building count.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗