Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “climate data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Enhancing Discoverability and Management of Atmospheric Data at Scale: Solutions from the ARM Data Center

The Atmospheric Radiation Measurement (ARM) is a multi-laboratory and multi-institutional U.S. Department of Energy (DOE) Office of Science National User Facility. The ARM Data Center (ADC), located at Oak Ridge National Laboratory, collects, archives, and shares vast atmospheric data crucial for climate research. The ADC manages over 7 PB of data from 460 instruments worldwide, processing it into more than 11,000 diverse data products using the Network Common Data Form (NetCDF) for machine-independent accessibility. The primary challenge addressed in this paper is the efficient management and distribution of vast and diverse datasets essential for the climate research community, enhancing accessibility through advanced tools like Data Discovery. The ADC has developed advanced infrastructure and software architecture to handle the continuous influx of heterogeneous data to enhance data discoverability, resulting in increased scientific collaboration. In 2023, users from over 34 countries downloaded and utilized ARM data, resulting in 1,455 publications. The ADC’s efforts have significantly improved the discoverability and usability of atmospheric data, fostering extensive scientific research and collaboration. This paper details the solutions implemented by the ADC team for efficient data discovery and distribution, and it demonstrates ARM’s capability of staging processed data for scientific analysis.

Shah, Chirag [ORNL] (ORCID:0000000203145737)↗

Bias Correction and Statistical Downscaling of Future Solar Irradiance Projections Using the NSRDB

Assessing renewable energy resources under future climate scenarios has been highlighted to understand potential impacts of future climate change in renewable generation on the power sector. Climate model projection has been recognized by the renewable energy community as a useful data set to analyze the impacts of future climate change on renewable resources. However, future climate projections generated from general circulation models (GCMs) contain inherent biases that need to be corrected for accurate analysis of future projections of climate variables. In addition, the coarse spatiotemporal resolution of GCMs needs to be improved for regional climate studies. In this work, we develop statistical methods to downscale future projections of global horizontal irradiance (GHI) in a computationally efficient way. Our approach builds statistical downscaling models that correct bias of climate projection of GHI and downscale the future GHI projection from daily-scale to hourly-scale. The National Solar Radiation Database (NSRDB) is used to calibrate the statistical models and validate the downscaled GHI projections across the contiguous United State (CONUS). Preliminary results show that the statistical approach efficiently downscales climate projections of GHI with a nBIAS of 3%, nMAE of 34 % and nRMSE of 46% calculated against NSRDB for CONUS. This study describes the implemented methodology and initial results as well as future research to create high-resolution climate data sets for solar energy applications.

analytical models↗

ARM Data-Oriented Metrics and Diagnostics Package for Climate Model Evaluation

A Python-based metrics and diagnostics package is currently being developed by the U.S. Department of Energy (DOE) Atmospheric Radiation Measurement (ARM) Infrastructure Team at Lawrence Livermore National Laboratory (LLNL) to facilitate the use of long-term, high-frequency measurements from the ARM Facility in evaluating the regional climate simulation of clouds, radiation, and precipitation. This metrics and diagnostics package computes climatological means of targeted climate model simulation and generates tables and plots for comparing the model simulation with ARM observational data. The Coupled Model Intercomparison Project (CMIP) model data sets are also included in the package to enable model intercomparison as demonstrated in Zhang et al. (2017). The mean of the CMIP model can serve as a reference for individual models. Basic performance metrics are computed to measure the accuracy of mean state and variability of climate models. The evaluated physical quantities include cloud fraction, temperature, relative humidity, cloud liquid water path, total column water vapor, precipitation, sensible and latent heat fluxes, and radiative fluxes, with plan to extend to more fields, such as aerosol and microphysics properties. Process-oriented diagnostics focusing on individual cloud- and precipitation-related phenomena are also being developed for the evaluation and development of specific model physical parameterizations. The version 1.0 package is designed based on data collected at ARM’s Southern Great Plains (SGP) Research Facility, with the plan to extend to other ARM sites. The metrics and diagnostics package is currently built upon standard Python libraries and additional Python packages developed by DOE (such as CDMS and CDAT). The ARM metrics and diagnostic package is available publicly with the hope that it can serve as an easy entry point for climate modelers to compare their models with ARM data. In this report, we first present the input data, which constitutes the core content of the metrics and diagnostics package in section 2, and a user's guide documenting the workflow/structure of the version 1.0 codes, and including step-by-step instruction for running the package in section 3.

54 ENVIRONMENTAL SCIENCES↗

Exascale Computing and Data Handling: Challenges and Opportunities for Weather and Climate Prediction

The emergence of exascale computing and artificial intelligence offer tremendous potential to significantly advance Earth system prediction capabilities. However, enormous challenges must be overcome to adapt models and prediction systems to use these new technologies effectively. A 2022 WMO report on exascale computing recommends “urgency in dedicating efforts and attention to disruptions associated with evolving computing technologies that will be increasingly difficult to overcome, threatening continued advancements in weather and climate prediction capabilities.” Further, the explosive growth in data from observations, model and ensemble output, and postprocessing threatens to overwhelm the ability to deliver timely, accurate, and precise information needed for decision-making. Artificial intelligence (AI) offers untapped opportunities to alter how models are developed, observations are processed, and predictions are analyzed and extracted for decision-making. Given the extraordinarily high cost of computing, growing complexity of prediction systems, and increasingly unmanageable amount of data being produced and consumed, these challenges are rapidly becoming too large for any single institution or country to handle. This paper describes key technical and budgetary challenges, identifies gaps and ways to address them, and makes a number of recommendations.

Atmosphere↗

Computational Modeling of Atmospheric Processes at Texas Southern University

Texas Southern University (TSU) is strengthening its research program in atmospheric chemistry and physics with a climate science emphasis by leveraging partnerships with the U.S. Department of Energy’s Atmospheric Radiation Measurement (ARM) Facility, Brookhaven National Laboratory (BNL), and the Tracking Aerosol Convection Interactions ExpeRiment (TRACER). This RDPP-supported program focuses on secondary organic aerosols (SOAs) and reactive atmospheric species that influence cloud formation, precipitation processes, and radiative forcing. SOAs play a critical role in cloud microphysics and Earth’s energy balance, yet the chemical and physical mechanisms governing SOA–cloud interactions remain a significant source of uncertainty in predictive climate models. Through computational modeling, observational data analysis, and national laboratory collaboration, this program develops a skilled cohort of students trained in atmospheric science, environmental data analysis, and climate-relevant modeling. These research experiences build technical competencies that are transferable to careers in government laboratories, academia, and industry. By engaging students from historically underrepresented communities in high-impact climate research, TSU expands participation in the atmospheric sciences workforce while contributing meaningful scientific insights to DOE-supported ARM research activities. This partnership strengthens national capacity in climate science and supports the development of the next generation of atmospheric researchers.

54 ENVIRONMENTAL SCIENCES↗

Characterizing climate pathways using feature importance on echo state networks

The 2022 National Defense Strategy of the United States listed climate change as a serious threat to national security. Climate intervention methods, such as stratospheric aerosol injection, have been proposed as mitigation strategies, but the downstream effects of such actions on a complex climate system are not well understood. The development of algorithmic techniques for quantifying relationships between source and impact variables related to a climate event (i.e., a climate pathway) would help inform policy decisions. Data-driven deep learning models have become powerful tools for modeling highly nonlinear relationships and may provide a route to characterize climate variable relationships. In this paper, we explore the use of an echo state network (ESN) for characterizing climate pathways. ESNs are a computationally efficient neural network variation designed for temporal data, and recent work proposes ESNs as a useful tool for forecasting spatiotemporal climate data. However, ESNs are noninterpretable black-box models along with other neural networks. The lack of model transparency poses a hurdle for understanding variable relationships. We address this issue by developing feature importance methods for ESNs in the context of spatiotemporal data to quantify variable relationships captured by the model. We conduct a simulation study to assess and compare the feature importance techniques, and we demonstrate the approach on reanalysis climate data. In the climate application, we consider a time period that includes the 1991 volcanic eruption of Mount Pinatubo. This event was a significant stratospheric aerosol injection, which acts as a proxy for an anthropogenic stratospheric aerosol injection. Furthermore, we are able to use the proposed approach to characterize relationships between pathway variables associated with this event that agree with relationships previously identified by climate scientists.

black-box models↗

Data readiness pipeline patterns for scientific AI at scale: Insights from climate, fusion, life sciences, and materials

This article examines how data readiness for AI principles apply to large scientific datasets used to train foundation models. We analyze archetypal workflows across four representative domains—climate, nuclear fusion, life sciences, and materials—to identify common preprocessing patterns and domain‐specific constraints. We introduce a two‐dimensional readiness model that combines canonical preprocessing patterns with a five‐level operational readiness scale, both tailored to high‐performance computing (HPC) environments. This construct helps outline key challenges in transforming large‐scale scientific data into formats suitable for scalable AI training. Together, these dimensions form a conceptual maturity matrix that characterizes scientific data readiness and guides infrastructure development toward standardized, cross‐domain support for scalable and reproducible AI for science. Finally, we evaluate this maturity matrix in the context of case studies including ClimaX (climate), AFLOW (materials), OpenFold (proteomics), and DIII‐D fusion disruption‐prediction workflows, from which we distill lessons learned and provide recommendations to guide practitioners in developing robust AI‐readiness pipelines. Finally, we discuss remaining cross‐cutting challenges that persist across scientific domains.

97 MATHEMATICS AND COMPUTING↗

Using feature importance as an exploratory data analysis tool on Earth system models

Abstract. Machine learning (ML) models are commonly used to generate predictions, but these models can also support the discovery of new science. Generating accurate predictions necessitates that a model captures the structure of the underlying data. If the structure is properly extracted, ML could be a useful exploratory and evidential tool. In this paper, we present a case study that demonstrates the use of ML for exploratory data analysis (EDA) in the climate space. We apply the ML explainability method of spatiotemporal zeroed feature importance (stZFI) to understand how climate-variable associations evolve over space and time. Our analyses focus on data from ensembles of Earth system models (ESMs) which provide data on different climate states and conditions. We elect to work with ESM ensembles since they allow us to compare feature importance across alternative scenarios not available with observed data. The ensembles also account for natural variability so that we can distinguish between signal and noise due to natural climate variability when computing feature importance. The use of perturbed initial condition ensembles introduces variability mimicking the natural variability in the atmosphere; thus the signals emerging using feature importance (FI) can be evaluated against the natural variability in the climate system. For our analyses, we consider the 1991 volcanic eruption of Mount Pinatubo, which was a large stratospheric aerosol injection. We explore the climate pathway associated with the eruption from aerosols to radiation to temperature at both the near-surface and stratospheric levels. In addition to applying the method to data generated from two different ESMs, we apply stZFI to reanalysis data to compare the associations identified by stZFI. We show how stZFI tracks the importance of aerosol optical depth over time on forecasting temperatures. This case study illustrates usefulness of an ML tool (stZFI) for EDA on a well-studied climate exemplar.

Ries, Daniel (ORCID:0000000250294647)↗

Application of the AI2 Climate Emulator to E3SMv2's Global Atmosphere Model, With a Focus on Precipitation Fidelity

Abstract Can the current successes of global machine learning‐based weather simulators be generalized beyond 2‐week forecasts to stable and accurate multiyear runs? The recently developed AI2 Climate Emulator (ACE) suggests this is feasible, based upon 10‐year simulations with a network trained on output from a physics‐based global atmosphere model using a grid spacing of approximately 110 km and forced by a repeating annual cycle of sea‐surface temperature. Here we show that ACE, without modification, can be trained to emulate another major atmospheric model, EAMv2, run at a comparable grid spacing for at least 10 years with similarly small climate biases—a prerequisite to wider applicability. With an analysis that combines multiple temporal, spatial, and frequency domain perspectives, we show that ACE faithfully represents the spatiotemporal structure of EAMv2 precipitation and related variables. Finally, we show that a pretrained ACE network is able to adapt to a new global climate model simulation data set with 10 fewer training steps than when starting from random initialization, all while still maintaining low levels of climate bias. Further analysis of these fine‐tuning experiments reveal ACE's intriguing ability to interpolate between distinct global climate models.

Duncan, James P. C.↗

National Climate Database (NCDB)

The National Climate Database (NCDB) is a high resolution, bias-corrected climate dataset consisting of the three most widely used variables of solar radiation- global horizontal (GHI), direct normal (DNI), and diffuse horizontal irradiance (DHI)- as well as other meteorological data. The goal of the NCDB is to provide unbiased high temporal and spatial resolution climate data needed for renewable energy modeling. The NCDB is modeled using a statistical downscaling approach with Regional Climate Model (RCM)-based climate projections obtained from the North American Coordinated Regional Climate Downscaling Experiment (NA-CORDEX; linked below). Daily climate projections simulated by the Canadian Regional Climate Model 4 (CanRCM4) forced by the second-generation Canadian Earth System Model (CanESM2) for two Representative Concentration Pathways (RCP4.5 or moderate emissions scenario and RCP8.5 or highest baseline emission scenario) are selected as inputs to the statistical downscaling models. The National Solar Radiation Database (NSRDB) is used to build and calibrate statistical models.

Array↗

Lack of clear standards and usable comparisons of downscaled climate projections pose a roadblock for US climate discovery and adaptation

Abstract The release of global climate projections coupled with the demand for local-resolution climate-forced meteorology has prompted many research groups to downscale these projections using various statistical, dynamical, and current machine learning techniques. Such downscaled datasets are being used to plan infrastructure and other community needs over the coming decades. Faced with roughly a dozen available US downscaled datasets, many practitioners ask, ‘What are the relevant differences between datasets?’ This work highlights the difficulty of comparing downscaled datasets and illustrates ways in which datasets differ even when using identical climate model input data. We show that substantial variability in precipitation projections arises from downscaling alone and that the downscaled dataset agreement varies depending on global climate projection. This analysis emphasizes the need for greater coordination and movement toward rigorous benchmarking of downscaling strategies within the downscaling research community, à la the land-modeling community, to better quantify downscaling dataset differences, strengths, and weaknesses for practitioners.

Hartke, Samantha H. (ORCID:0000000202394723)↗

Data and scripts associated with “Non-random processes impacting organic matter chemistry are maximized in mid-order streams”

NOTE: The manuscript associated with this data package is currently in review. The data may be revised based on reviewer feedback. Upon manuscript acceptance, this data package will be updated with the final dataset and additional metadata. This data package is associated with the publication “Non-random processes impacting organic matter chemistry are maximized in mid-order streams” submitted to Limnology and Oceanography (L&O) by Danczak et al. (in review). This package contains data and scripts used to investigate dissolved organic matter (DOM) molecular chemistry and diversification processes across 47 surface-water sampling sites in the Yakima River Basin, Washington, USA, during an August 2021 sampling campaign. The package contains analyses of ultrahigh-resolution Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS), geochemical measurements, geospatial attributes, molecular diversity, and meta-metabolome ecological null models needed to reproduce the main manuscript results. The underlying field data were pulled from exising data packages at https://doi.org/10.15485/1892052 (Fulton et al., 2022) and https://doi.org/10.15485/1898914 (Grieger et al., 2022). For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. We thank the following organizations for providing access to field locations for sample collection: the United States Forest Service, Washington Department of Fish and Wildlife, Washington Department of Natural Resources, the Confederated Tribes and Bands of the Yakama Nation, and the Cowiche Canyon Conservatory. Research was conducted under Washington State Parks and Recreation Commission Scientific Research Permit #210901. We are grateful to the Yakama Nation Tribal Council and Yakama Nation Fisheries for their collaboration in facilitating sample collection and ensuring data usage aligns with their values and worldview. This data package contains an R-Markdown file for analyses and five folders: (1) Data, (2) Geospatial Data, (3) Supplemental_Files, (5) Figures_pdf, (4) and src. The Data folder contains tabular inputs and derived files used in the manuscript analysis. The Geospatial Data folder contains climate and water-balance, hydrologic, land-cover, population/regional water-use, stream, topographic, and stream-order attribute CSV files. The src folder contains scripts used to process data, run analyses, and generate figures. The Figures_pdf folder contains manuscript figure outputs. The Supplemental_Files folder contains supplemental analysis products. All files are .csv, .pdf, .html, .png, .R, .Rmd, .svg, or .tre. This data package is associated with the rcfsa-RC2-SPS_Null_Modeling repository found at https://github.com/river-corridors-sfa/rcfsa-RC2-SPS_Null_Modeling.

54 ENVIRONMENTAL SCIENCES↗

Informing forest carbon inventories under the Paris Agreement using ground-based forest monitoring data

Human interactions with forests have shaped Earth's climate for millennia and will continue to do so as we target net-zero emission goals. Accurately characterizing these climate impacts requires making reliable forest carbon data available for forest monitoring and planning. Here, we develop a semi-automated process for submitting forest carbon measurements from the largest relevant scientific database to the International Panel on Climate Change's Emission Factor Database, which currently has sparse forest carbon data. Building this bridge from scientific research to international policy is an important step towards managing forests in a net-zero motivated future. Humans have been influencing Earth's climate via transformative impacts on forests for millennia, and forests are now recognized as critical to climate change mitigation under the Paris Agreement. The efficacy of climate change mitigation planning and reporting depends on quality data on forest carbon (C) stocks and changes. The Emission Factor Database (EFDB) of the International Panel on Climate Change (IPCC) is intended to be a definitive source for such data, but needs comprehensive and well-documented data to be so. To facilitate submission of forest C estimates from scientific studies to EFDB, we develop and document a process for semi-automated data submission from the Global Forest C database (ForC v4.0), which is the largest compilation of ground-based forest C estimates. We then assess the data currently available through ForC and provide recommendations for improving forest data collection, analysis, and reporting. As of September 2024, ForC contained ~19,286 records potentially relevant to EFDB, 1068 of which had been submitted and posted to EFDB. These represented 19% of the total EFDB records for forest land. Records were unevenly distributed across variables and geographic regions. ForC records (37%) reviewed could not be submitted because the original publication lacked required information. In the future, ground-based forest C estimates should target gaps in the record, and studies should ensure that they report all information necessary for inclusion in EFDB. Given that climate change is rapidly impacting the world's forests, timely reporting of recent estimates will be critical to accurate forest C inventories.

54 ENVIRONMENTAL SCIENCES↗

Assessing the performance of solar radiation management geoengineering simulations

Offsetting the global warming caused by anthropogenic increases in atmospheric greenhouse gases by deliberate injection of aerosols into the stratosphere is the most studied of solar radiation management geoengineering schemes. The long-term success or failure of such schemes in achieving their stated goals is assessed by comparing simulated geoengineered temperature, precipitation and tropical cyclones metrics to equivalent fields in the simulated targeted climate simulations. Results using available data sets from three single model stabilized climate target experiments and three multimodel climate change reduction experiments are presented and compared against a measure of internal variability. While all but one experimental scheme is successful in achieving their targeted global mean annual surface temperature, their success at regional scales varies significantly and is often larger than the internal variability metric used here.

climate model evaluation↗

Legacy Effects of Cropping System and Precipitation Influence the Core Camelina sativa Microbiome

Camelina ( Camelina sativa L.) is a potential biofuel crop and beneficial rotation crop in dryland cropping systems. Little is known about camelina microbiota or the legacy effect of soil origin/cropping system zones on camelina-associated microbiome assembly. To explore camelina-microbe associations, we grew camelina in the greenhouse using soil transplanted from 33 locations in the dryland wheat production area of eastern Washington. Bacterial, archaeal, and fungal communities from bulk soil, rhizosphere, and endosphere were characterized with 16S rRNA and internal transcribed spacer amplicon sequencing and were analyzed alongside site-specific climatic and edaphic data. We found that soil from the highest precipitation zone had higher alpha diversity than soil from the driest zone, but this effect was not seen in the greenhouse rhizosphere or endosphere. Plant compartment, cropping system zone, and soil origin all significantly influenced microbial composition, with soil pH and organic matter, as well as precipitation at origin, as major predictors. Analysis of abundance–occupancy distributions showed that the Actinobacteriota Aeromicrobium and Marmoricola and the fungus Pseudogymnoascus in the rhizosphere were plant-selected, while the endosphere was characterized by a number of Actinobacteriota, Rhizobium, and Clostridium. Sphingomonas amplicon sequence variants were also consistently enriched in the rhizosphere, suggesting that they are present in soils collected throughout eastern Washington and may represent good candidate biostimulants. Several lignin decomposing fungi had site-specific rhizospheric distributions, suggesting that they may be dispersal-limited or result from the legacy effect of long-term wheat cropping. Overall, this study contributes to our understanding of microbiome assembly in and on camelina roots while also highlighting the potential impact of cropping history on soil- and plant-associated microbiomes. [Formula: see text] The author(s) have dedicated the work to the public domain under the Creative Commons CC0 “No Rights Reserved” license by waiving all of his or her rights to the work worldwide under copyright law, including all related and neighboring rights, to the extent allowed by law, 2025.

Barnes, Elle M↗

BSDF Data generation for daylight applications: A call for international standardization

Standardized methods for generating angle-dependent, bidirectional, solar-optical properties for complex fenestration systems do not exist, which means that energy and daylight evaluations in building performance simulations often suffer from major inaccuracies. This position paper provides an overview of state-of-the-art data-driven methods for characterizing light scattering properties of fenestration materials and blind systems (e.g. fabrics, metal slats, patterned glazing), validation via laboratory, simulation and field tests, and salient issues in support of standardization of such methods via the International Standardization Organization (ISO). The ISO standard is intended to provide the fundamental underpinnings for recently mandated daylight standards that rely on bidirectional scattering distribution function data for climate-based daylight modelling and building performance simulations.

Geisler-Moroder, D.↗

2024 IUFRO Tree Biotechnology Conference (Aug 4-8, 2024)

The 2024 IUFRO Tree Biotechnology Conference is the biennial meeting on genomics, molecular biology, and biotechnology of forest trees, associated with the IUFRO Working Party 2.04.06. This year's meeting was held in Annapolis, MD, USA from August 4th to 8th and was hosted by Yiping Qi (University of Maryland), Edward Eisenstein (University of Maryland), Gary Coleman (University of Maryland), and Heather Coleman (Syracuse University). The conference covered seven topics over the course of five days: 1) Biological and ecological insights from OMICS, 2) Advancing technologies for targeted trait manipulation and acceptability to diverse tree species, 3) Genes, development, and physiology, 4) Translating genomics and biotechnology to practice, 5) Trees in a changing world, 6) Genetic and phenotypic diversity for breeding and genomic selection, and 7) Biotechnology for biomaterials and bioeconomy. In addition to the sessions, there were two plenary sessions, provided by John Ralph (University of Wisconsin) and Tanja Pyrhäjärvi (University of Helsinki). The meeting celebrated the second awardees of the newly created IUFRO WG 2.04.06 Award: Excellence in Forest Molecular Biology and Genomics, which was presented to Chung-Jui (C.J.) Tsai (University of Georgia). Greg Goralogia (Oregon State University) was the recipient of the associated Early Career Award. The scientific presentations at the conference highlighted cutting-edge advancements in many facets of forest biotechnology research, including applications of genomic selection in forest genetics and breeding, the use of genetic editing, tree physiology, stress response, molecular breeding, wood development, "omics" technologies, and the social and economic impacts of genetically modified (GM) trees. Scientific take homes from the meeting include the power of NMR to dissect the composition of lignin, the genomic diversity of forest trees that has enormous potential for tree improvement and the integration of systems biology with climate and geographical data. The conference attracted a mix of students (25), postdoctoral fellows (32), and scientists from academia (66) and industry (18). In all, the conference was attended by 141 registered participants, representing 20 countries that participated in 23 invited lectures (including 6 'early-career' keynotes), 27 voluntary talks and 61 poster presentations. Support for the conference was drawn from a wide variety of Academia, Industry, and Government sources, and included financial support from several tree improvement companies. Overall, the conference was a great success, providing an exceptional mix of science and social activities in a relaxed and collegial atmosphere. More information about the meeting can be found at treebiotech.org. The next meeting will be held in Stellenbosch, South Africa, in 2026, hosted jointly by Zander Myburg, Dave Drew (University of Stellenbosch,) and Sanushka Naidoo (University of Pretoria, FABI).

59 BASIC BIOLOGICAL SCIENCES↗