Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “research data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

Using probability distribution function as a scaling approach to incorporate soil heterogeneity into biogeochemical models for greenhouse gas predictions (Final Technical Report)

The project investigated biogeochemical processes at terrestrial-aquatic interfaces (TAIs), focusing on soil microsite heterogeneity and its impact on greenhouse gas (GHG) fluxes. Using laboratory experiments, modeling, and data integration, researchers explored redox-driven microbial processes under fluctuating hydrological conditions. Key advancements included modifying the DAMM-GHG model to incorporateelectron acceptor availability and enhancing the AquaMEND model for improved microbial metabolism representation. Results highlighted microsite redox variability as a key driver of GHG fluxes, informing Earth system models. The project fostered interdisciplinary collaborations, student training, and the development of novel modeling frameworks to improve Earth'senergy budget.

54 ENVIRONMENTAL SCIENCES↗

Data Cards for Standardized Metadata Across DOE-Aligned Data Initiatives: Toward Transparent, Interoperable, and Governed Dataset Documentation

As data-intensive research, advanced computing, and artificial intelligence become increasingly central to scientific and operational workflows, the need for consistent, transparent, and machine-actionable documentation has grown correspondingly. Multiple DOE-aligned communities—including Office of Science, Genesis Mission, American Science Cloud (AmSC), National Nuclear Security Administration (NNSA) stewardship and governance, and related cross-laboratory collaborations—have independently developed metadata practices to support discovery, access, reuse, repository deposit, and compliance.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Report of the 2026 Workshop on Next-Generation Ecosystems for Scientific Computing: Harnessing Community, Software, and AI for Cross-Disciplinary Team Science

Scientific computing is undergoing rapid transformation as advances in artificial intelligence, heterogeneous computing, automation, and data-intensive research reshape not only computational tools but also the institutions, workforce models, and collaborative practices that support scientific discovery. This report synthesizes insights from the 2026 Workshop on Next-Generation Ecosystems for Scientific Computing, the second in a three-year series focused on strengthening scientific computing ecosystems through socio-technical co-design. Workshop discussions identified four interdependent strategic themes: software ecosystems for AI-enabled scientific discovery; trust, validation, and traceability; human-AI teaming and paradigm shifts; and workforce, pedagogy, and governance. The report translates these themes into eight priorities for community action spanning shared research infrastructure, trust and traceability, user experience, human-AI teaming, workforce development, cross-sector coordination, stewardship and sustainability, and evaluation of scientific value. Together, these priorities outline directions for building scientific computing ecosystems that remain trustworthy, sustainable, innovative, and resilient as AI assumes a growing role in scientific work.

AI↗

Bioinformatic teaching resources - for educators, by educators - using KBase, a free, user-friendly, open source platform

Over the past year, biology educators and staff at the Department of Energy Systems Biology Knowledgebase (KBase) initiated a collaborative effort to develop a curriculum for bioinformatics education. KBase is a free and easily accessible data science platform that integrates many bioinformatics resources into a graphical user interface built upon reproducible analysis notebooks. KBase held conversations with college and high school instructors to understand how KBase could potentially support their educational goals. These conversations morphed into a working group of biological and data science instructors that adapted the KBase platform to their curriculum needs, specifically around concepts in Genomics, Metagenomics, Pangenomics, and Phylogenetics. The KBase Educators Working Group developed modular, adaptable, and customizable instructional units. Each instructional module contains teaching resources, publicly available data, analysis tools, and markdown capability to tailor instructions and learning goals for each class. The online user interface enables students to conduct hands-on data science research and analyses without requiring programming skills or their own computational resources (these are provided by KBase). Alongside these resources, KBase continues to work with instructors, supporting the development of additional curriculum modules. For anyone new to the platform, KBase, and the growing KBase Educators Organization, provides a community network, accompanied by community-sourced guidelines, instructional templates, and peer support to use KBase within a classroom whether virtual or in-person.

59 BASIC BIOLOGICAL SCIENCES↗

Using automated machine learning for the upscaling of gross primary productivity

Estimating gross primary productivity (GPP) over space and time is fundamental for understanding the response of the terrestrial biosphere to climate change. Eddy covariance flux towers provide in situ estimates of GPP at the ecosystem scale, but their sparse geographical distribution limits larger-scale inference. Machine learning (ML) techniques have been used to address this problem by extrapolating local GPP measurements over space using satellite remote sensing data. However, the accuracy of the regression model can be affected by uncertainties introduced by model selection, parameterization, and choice of explanatory features, among others. Recent advances in automated ML (AutoML) provide a novel automated way to select and synthesize different ML models. In this work, we explore the potential of AutoML by training three major AutoML frameworks on eddy covariance measurements of GPP at 243 globally distributed sites. We compared their ability to predict GPP and its spatial and temporal variability based on different sets of remote sensing explanatory variables. Explanatory variables from only Moderate Resolution Imaging Spectroradiometer (MODIS) surface reflectance data and photosynthetically active radiation explained over 70 % of the monthly variability in GPP, while satellite-derived proxies for canopy structure, photosynthetic activity, environmental stressors, and meteorological variables from reanalysis (ERA5-Land) further improved the frameworks' predictive ability. We found that the AutoML framework Auto-sklearn consistently outperformed other AutoML frameworks as well as a classical random forest regressor in predicting GPP but with small performance differences, reaching an r 2 of up to 0.75. We deployed the best-performing framework to generate global wall-to-wall maps highlighting GPP patterns in good agreement with satellite-derived reference data. This research benchmarks the application of AutoML in GPP estimation and assesses its potential and limitations in quantifying global photosynthetic activity.

54 ENVIRONMENTAL SCIENCES↗

Merged Observatory Data Files (MODFs): an integrated observational data product supporting process-oriented investigations and diagnostics

A large and ever-growing body of geophysical information is measured in campaigns and at specialized observatories as a part of scientific expeditions and experiments. These collections of observed data include many essential climate variables (as defined by the Global Climate Observing System) but are often distinguished by a wide range of additional non-routine measurements that are designed to not only document the state of the environment but also the drivers that contribute to that state. These field data are used not only to further understand environmental processes through observation-based studies but also to provide baseline data to test model performance and to codify understanding to improve predictive capabilities. To address the considerable barriers and difficulty in utilizing these diverse and complex data for observation–model research, the Merged Observatory Data File (MODF) concept has been developed. A MODF combines measurements from multiple instruments into a single file that complies with well-established data format and metadata practices and has been designed to parallel the development of corresponding Merged Model Data Files (MMDFs). Using the MODF and MMDF protocols will facilitate the evolution of model intercomparison projects into model intercomparison and improvement projects by putting observation and model data “on the same page” in a timely manner. The MODF concept was developed especially for weather forecast model studies in the Arctic. The surprisingly complex process of implementing MODFs in that context refined the concept itself. Thus, this article explains the concept of MODFs by providing details on the issues that were revealed and resolved during that first specific implementation. Detailed instructions are provided on how to make MODFs, and this article can be considered a MODF creation manual.

54 ENVIRONMENTAL SCIENCES↗

Mechanical and durability properties of ultra-high-performance concrete of spent nuclear fuel dry storage systems: a review

Dry storage systems are used for interim storage of spent nuclear fuel (SNF). However, with the growing need to extend the operational periods of these systems, there are concerns about the degradation of their concrete overpacks, which could compromise the system's structural integrity and safety during hazardous events. Traditional concrete mixtures used in SNF dry storage systems have remained largely unchanged since their inception and often use conventional ingredients. These materials are susceptible to degradation mechanisms such as chemical attacks, alkali-silica reactions (ASR), and freeze–thaw cycles, which can lead to a loss of strength and durability over time. To address these challenges, this paper reviews the application of ultra-high-performance concrete (UHPC) as a promising alternative for spent nuclear fuel dry storage system overpacks. UHPC offers superior mechanical properties, exceptional durability, and reduced susceptibility to degradation mechanisms compared to conventional concrete. This paper focuses on the role of supplementary cementitious materials (SCMs) such as silica fume, fly ash, and metakaolin in enhancing UHPC performance for SNF storage applications. These SCMs have been shown to significantly improve the material’s microstructure, strength, and resistance to environmental stressors typically encountered in SNF storage environments. Moreover, incorporating SCMs supports sustainable construction by reducing cement consumption and associated carbon emissions. The review brings together existing research and experimental data, providing insights for engineers and researchers on developing UHPC mixtures that meet the rigorous demands of spent nuclear fuel dry storage systems, extending their service life and minimizing inspection intervals.

36 - MATERIALS SCIENCE↗

Sustainable Port Operations: Powered by NREL

Seaports are vital economic hubs that allow the United States to compete on a global scale. But the heavy vehicles and cargo equipment that enable their operations also emit harmful air pollutants and greenhouse gas emissions. For nearly two decades, National Renewable Energy Laboratory (NREL) researchers have worked toward comprehensive seaport decarbonization. They fuse world-class analysis with deep vehicle and transportation systems knowledge to guide strategic deployment of low- and zero-emissions vehicles, charging and refueling infrastructure, and grid improvements. Together, these capabilities can enable sustainable port operations. This fact sheet outlines major seaport and airport decarbonization capabilities across the laboratory, including: fleet research, energy data, and insights for decarbonization; comprehensive hydrogen infrastructure deployment; optimized charging through grid integration; strategic blueprinting for clean, optimized technology deployment; and integrating diversity, equity, inclusion, and accessibility considerations into decarbonization efforts.

ADVANCED PROPULSION SYSTEMS,ENERGY CONSERVATION, C↗

GeoRePORT (Geothermal Resource Portfolio Optimization & Reporting Technique)

The Geothermal Resource Portfolio Optimization and Reporting Technique (GeoRePORT) Protocol provides a system for reporting resource grade and project readiness level. It is particularly useful for describing early-stage exploration projects. GeoRePORT can assist in evaluating project risk and return, identifying gaps in reported data, evaluating research and design impacts, and gathering insights on successes and failures. It helps users objectively and quantitatively compare project potential in geological, technical, and socioeconomic areas.

Young, Katherine↗

Geospatial Data Platform for All

Spatiotemporal data has evolved in scale due to augmented use in cross-domain applications. Simultaneously, there is substantial growth in the availability of Geographic Information Systems (GIS) data provided by the United States Geological Survey (USGS) along with other federal, state, county, or local agencies through open-data portals and public access APIs. However, data availability does not equate with accessibility. Large-scale analyses and applications require robust, performant data management with co-location of data storage and computing. The insufficiency of data management infrastructure compels researchers to adopt ad hoc project- specific GIS data storage solutions (e.g., copying data to High-Performance computer file systems). As an ad hoc storage strategy does not scale, it hampers cross-domain analyses causing difficulty in data reuse and utilizing existing code bases. Furthermore, GIS data is complex and requires expertise to analyze and manipulate due to its intricate data structures and data-specific projection transformations. Despite the challenges, we recognize that derived GIS data products, e.g., satellite or LIDAR-based images, can be used in downstream applications such as AI by domain, but non-GIS experts. To address the data needs and overcome the challenges, we are working towards a GIS Data Platform focused on efficient data storage, data discovery and access, and an API to enable common workflows. We propose a knowledge-graph (KG) approach for data discovery, whereby datasets are semantically linked to higher- level constructs such as projects and research areas. The semantic data links enable researchers to explore datasets in a top-down approach by specifying relevant and meaningful terms (assists in finding hidden data). An advantage is that the nodes and edges in a knowledge graph create built-in semantic documentation. Deeper spatiotemporal connections between data sources can be encoded via Graph Neural Networks (GNN) (Zhang et al., 2021). The KG approach can be extended to integrate the data itself in a Virtual KG (VKG). Our work will derive inspiration from large-scale VKG efforts that have been undertaken or are currently underway as part of the OpenStreetMap project (Ding et al., 2021). For DOE Data Days, we share the proposed geospatial data platform hybrid (cloud/on-prem) architecture, our work-to-date on storing, retrieving, and transforming LiDAR and raster data relevant to two important NREL use-cases, including the Renewable Energy Potential (reV) Model, and present our proposal for a KG based data discovery engine.

data platform↗

Developing an Interactive Landscape for Mobility Resources: Preprint

As the world continues to be increasingly driven by data, the ways researchers and professionals sort and collect this data is critical. In the world of mobility data, new levels of data from public transportation systems, location services, and other means are being lost due to how little organization exists. Much of the data is proprietary, and there are few if any de jure or even de facto standards connecting data. There is also little knowledge about the gaps that exist in the data. In this project, we created an interactive landscape where mobility resources are categorized and organized in an easy to use, living document. We made this landscape with open-source code from the CNCF Cloud Native Landscape and repurposed it to the mobility data's needs. Additionally, unlike previous sources that organize mobility data, this document can be updated through GitHub by those in the field to keep its sources relevant. Following the creation of a beta version of the landscape, we conducted several interviews with industry researchers and professionals to ensure the landscape would be useful. The result is an online hub where mobility researchers and resource creators can easily access research and collaborate.

ADVANCED PROPULSION SYSTEMS↗

A Unified Data Infrastructure for Biological and Environmental Research: A Report from the BER Advisory Committee

The Biological and Environmental Research (BER) program within the U.S. Department of Energy (DOE) Office of Science supports large-scale data generation efforts across its two divisions: Biological Systems Science and Earth and Environmental Systems Sciences. These efforts include user facilities in atmospheric radiation measurements, genomics, metabolomics, proteomics, compute, and imaging. In addition, BER supports the development of plant-based fuels; research in biosystems design, environmental microbiomes, and atmospheric systems; energy flux monitoring; climate-based ecosystem experiments; pathogen biopreparedness; and modeling of climate, urban interfaces, and interactions between people and energy resources. For data access, BER supports community data services at its user facilities, along with specialized data initiatives for Earth and environmental science, climate modeling, genomic and microbial analysis, and multisector dynamics modeling.

54 ENVIRONMENTAL SCIENCES↗

32 examples of LLM applications in materials science and chemistry: towards automation, assistants, agents, and accelerated scientific discovery

Abstract Large language models (LLMs) are reshaping many aspects of materials science and chemistry research, enabling advances in molecular property prediction, materials design, scientific automation, knowledge extraction, and more. Recent developments demonstrate that the latest class of models are able to integrate structured and unstructured data, assist in hypothesis generation, and streamline research workflows. To explore the frontier of LLM capabilities across the research lifecycle, we review applications of LLMs through 32 total projects developed during the second annual LLM hackathon for applications in materials science and chemistry, a global hybrid event. These projects spanned seven key research areas: (1) molecular and material property prediction, (2) molecular and material design, (3) automation and novel interfaces, (4) scientific communication and education, (5) research data management and automation, (6) hypothesis generation and evaluation, and (7) knowledge extraction and reasoning from the scientific literature. Collectively, these applications illustrate how LLMs serve as versatile predictive models, platforms for rapid prototyping of domain-specific tools, and much more. In particular, improvements in both open source and proprietary LLM performance through the addition of reasoning, additional training data, and new techniques have expanded effectiveness, particularly in low-data environments and interdisciplinary research. As LLMs continue to improve, their integration into scientific workflows presents both new opportunities and new challenges, requiring ongoing exploration, continued refinement, and further research to address reliability, interpretability, and reproducibility.

Computer Science↗

Automated Shift Detection in Sensor-Based PV Power and Irradiance Time Series

PV power and irradiance sensor-based measurements are prone to error, resulting in issues such as time series data shifts. In this research, a changepoint detection (CPD) algorithm that automatically detects data shifts in sensor-based time series is introduced. Data shift periods in 101 daily PV power and irradiance time series were labeled manually by two solar experts. These data streams represent sensor-based measurements, and display a variety of data shift behaviors. A changepoint detection algorithm was tuned using the 101 labeled data streams, with each model configuration's ability to detect labeled changepoints benchmarked using metrics such as F1-score, recall, and Rand Index. Best performing models on seasonality-corrected data streams include the Pruned Exact Linear (PELT) method, the Binary Segmentation method, and the Bottom-Up method, all scoring an average F1-score of 0.76 or greater at detecting labeled changepoints within a 30-day window across the labeled data sets. Pending approval, we plan to release the labeled data sets for this research on NREL's DuraMAT Data Hub, and the associated algorithm in the Python PVAnalytics package. By supplying the training sets and algorithm, we hope to encourage further development in this research space.

data shift↗

A three-year dataset supporting research on building energy management and occupancy analytics

Abstract This paper presents the curation of a monitored dataset from an office building constructed in 2015 in Berkeley, California. The dataset includes whole-building and end-use energy consumption, HVAC system operating conditions, indoor and outdoor environmental parameters, as well as occupant counts. The data were collected during a period of three years from more than 300 sensors and meters on two office floors (each 2,325 m 2 ) of the building. A three-step data curation strategy is applied to transform the raw data into research-grade data: (1) cleaning the raw data to detect and adjust the outlier values and fill the data gaps; (2) creating the metadata model of the building systems and data points using the Brick schema; and (3) representing the metadata of the dataset using a semantic JSON schema. This dataset can be used in various applications—building energy benchmarking, load shape analysis, energy prediction, occupancy prediction and analytics, and HVAC controls—to improve the understanding and efficiency of building operations for reducing energy use, energy costs, and carbon emissions.

24 POWER TRANSMISSION AND DISTRIBUTION↗

BioHackathon 2015: Semantics of data for life sciences and reproducible research

We report on the activities of the 2015 edition of the BioHackathon, an annual event that brings together researchers and developers from around the world to develop tools and technologies that promote the reusability of biological data. We discuss issues surrounding the representation, publication, integration, mining and reuse of biological data and metadata across a wide range of biomedical data types of relevance for the life sciences, including chemistry, genotypes and phenotypes, orthology and phylogeny, proteomics, genomics, glycomics, and metabolomics. We describe our progress to address ongoing challenges to the reusability and reproducibility of research results, and identify outstanding issues that continue to impede the progress of bioinformatics research. We share our perspective on the state of the art, continued challenges, and goals for future research and development for the life sciences Semantic Web.

97 MATHEMATICS AND COMPUTING↗