Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Large Dataset Processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Fostering Geothermal Machine Learning Success: Elevating Big Data Accessibility and Automated Data Standardization in the Geothermal Data Repository

The Department of Energy's (DOE's) Geothermal Data Repository (GDR) has implemented improvements to both its data lakes and its data standards and automated data pipelines. The GDR data lakes have reduced storage and compute-related barriers to using large geothermal datasets, enabling these large datasets to be accessed by anyone with a modern computer and internet access. More recently, the GDR has been working to further reduce barriers through streamlining the data intake process, educating users on the process and requirements, and helping users access data from the data lakes. These improvements have augmented the quantity of datasets the GDR is able to accept into its data lakes and have enabled users who are new to cloud tools to access these datasets more easily, overall increasing the accessibility of big geothermal data for use in machine learning and other projects. In addition, the GDR now has built-in data standards and pipelines for drilling data, geospatial data, and distributed acoustic sensing (DAS) data. These standardization efforts aim to enhance the real-world applicability of geothermal machine learning outcomes by improving the quality of training data. Specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, thus allowing more time for actual research. By automating this process, the burden of standardization is lifted from the user, ultimately increasing the availability of standardized data. This paper provides an update on recent improvements made to the GDR's data lakes and automated data pipelines, including: (1) streamlining the data lake intake process, (2) better educating users on the process and requirements through a new data lakes page, (3) adding data lake direct access links to GDR data lake submission pages, (4) implementing a DAS data pipeline to convert DAS data uploaded in SEG-Y format to a standardized hierarchical data format v5 (HDF5), (5) extending this pipeline to encompass data in the GDR data lake, (6) adding metadata requirements for geospatial data, (7) making user interface/user experience (UX) enhancements to the data pipelines' documentation pages, and (8) improving the GDR's data standards and pipelines pages to better guide users in ensuring that their data is standardized by the GDR's automated data pipelines. 2024 Geothermal Resources Council. All rights reserved.

accessibility↗

Improving microstructures segmentation via pretraining with synthetic data

Image analysis of material microstructures through microscopy is an integral capability in the field of materials science. The topological and chemical information obtained through microscopy allow us to draw vital connections between material microstructures, properties, and processing. While scanning electron microscopy (SEM) is able to yield a considerable wealth of information interpretable by the intuition of experts, there has been considerable interest in using machine learning, convolutional neural networks (CNNs) in particular, for such image analysis task. Training CNNs for an image analysis task requires a large annotated dataset. However, in many materials science applications, obtaining a large annotated dataset is cost and labor intensive. In this work, we study the use of synthetic data to enlarge the available annotated experimental data of uranium oxide. We utilize a modified Potts model to simulate uranium oxide particles with morphologies similar to those observed experimentally. We then leverage an image-to-image translation model to synthesize the simulated particles as if they are acquired with SEM. Through this process, we obtain pairs of particle images and their corresponding SEM representations, which corresponds to pairs of annotations and images. Unlike previous works, we leverage synthetic data for pretraining a CNN model prior, and finetune that model further with experimental data. We experimentally demonstrate that using synthetic data as incremental learning process benefits the overall performance compared to training a model on combined synthetic and experimental data.

36 MATERIALS SCIENCE↗

AI-NERD: Elucidation of relaxation dynamics beyond equilibrium through AI-informed X-ray photon correlation spectroscopy

Abstract Understanding and interpreting dynamics of functional materials in situ is a grand challenge in physics and materials science due to the difficulty of experimentally probing materials at varied length and time scales. X-ray photon correlation spectroscopy (XPCS) is uniquely well-suited for characterizing materials dynamics over wide-ranging time scales. However, spatial and temporal heterogeneity in material behavior can make interpretation of experimental XPCS data difficult. In this work, we have developed an unsupervised deep learning (DL) framework for automated classification of relaxation dynamics from experimental data without requiring any prior physical knowledge of the system. We demonstrate how this method can be used to accelerate exploration of large datasets to identify samples of interest, and we apply this approach to directly correlate microscopic dynamics with macroscopic properties of a model system. Importantly, this DL framework is material and process agnostic, marking a concrete step towards autonomous materials discovery.

36 MATERIALS SCIENCE↗

Constrained GAN-Generated X-Ray CT Data For Self-Supervised And Foundation-Model Segmentation Of Concrete Microstructures

Three-dimensional characterization of materials using X-ray computed tomography (XCT) is challenging due to the complexity of internal structures, noise, and variations in resolution. Traditional computer vision models often struggle to accurately segment these images, particularly in domain-specific applications like materials science. While supervised deep learning approaches have been developed to address the limitations of conventional algorithms, they typically require large amounts of labeled training data and often fail to generalize across different datasets. Self-supervised, few-and zero-shot learning methods have gained prominence in natural image processing and segmentation tasks, but their application to scientific imaging remains limited due to the unique structural complexity, noise, and textural artifacts present in materials science data. In this work, we investigate how domain adaptation, leveraging physics-based and GAN-generated synthetic data, impacts segmentation performance. We introduce a modified Contrastive Unpaired Translation (CUT) model designed to generate realistic labeled data, which can be used for training, pre-training, and fine-tuning segmentation models for real XCT microstructure data. We evaluate the performance of two segmentation approaches: a self-supervised network (SSL-ALPNet) and a foundation model (Segment Anything Model), assessing their improvements when pre-trained and/or fine-tuned on the synthesized data. Our results demonstrate that leveraging synthetic data significantly enhances segmentation performance, particularly in challenging materials science applications.

Ziabari, Amir [ORNL] (ORCID:000000034776457X)↗

Interactive Visualization of Near Real Time and Production Global Precipitation Measurement (GPM) Mission Data Online Using CesiumJS

Advancements in the capabilities of JavaScript frameworks and web browsing technology make online visualization of large geospatial datasets viable. Commonly this is done using static image overlays, prerendered animations, or cumbersome geoservers. These methods can limit interactivity andor place a large burden on server-side post-processing and storage of data. Geospatial data, and satellite data specifically, benefit from being visualized both on and above a three-dimensional surface. The open-source JavaScript framework CesiumJS, developed by Analytical Graphics, Inc., leverages the WebGL protocol to do just that. It has entered the void left by the abandonment of the Google Earth Web API, and it serves as a capable and well-maintained platform upon which data can be displayed. This paper will describe the technology behind the two primary products developed as part of the NASA Precipitation Processing System STORM website: GPM Near Real Time Viewer (GPMNRTView) and STORM Virtual Globe (STORM VG). GPMNRTView reads small post-processed CZML files derived from various Level 1 through 3 near real-time products. For swath-based products, several brightness temperature channels or precipitation-related variables are available for animating in virtual real-time as the satellite-observed them on and above the Earths surface. With grid-based products, only precipitation rates are available, but the grid points are visualized in such a way that they can be interactively examined to explore raw values. STORM VG reads values directly off the HDF5 files, converting the information into JSON on the fly. All data points both on and above the surface can be examined here as well. Both the raw values and, if relevant, elevations are displayed. Surface and above-ground precipitation rates from select Level 2 and 3 products are shown. Examples from both products will be shown, including visuals from high impact events observed by GPM constellation satellites.

satellite precipitation measurement↗

Synergy of Satellite Radiation, Precipitation, and Other Meteorological Variable Observations for Global Mean Sea Surface Turbulent Heat Flux Estimation

Sea surface turbulent heat flux is one of the key components in global freshwater and energy balances and plays an important role in atmospheric dynamics, thermodynamics and general circulation. This turbulent heat flux is dominantly decided by sea surface latent heat release with some contribution from sensible heat exchange. The flux and its anomaly could significantly affect ocean heat storage and ocean circulation. They are part of and have great impacts on climate variability. Though the turbulent heat flux is extremely important for the climate and weather systems, only very limited ship and buoy observations of the flux are available over open oceans. There are no global, operational, direct measurements of this crucial meteorological variable. Global estimates are, basically, indirectly calculated from bulk turbulent flux parameterization with a combination of satellite column water vapor, sea surface water temperature, and wind speed observations. Critical parameters such as sea surface air temperature and humidity are estimated from empirical relations of theses variables with the water vapor and sea surface water temperature, respectively. Lacking accurate knowledge on surface air temperature and humidity, uncertainties in the parameterization and potential changes in the non-linear turbulent processes with long-term climate variations could cause large errors in estimated long-term turbulent fluxes from the indirect method as shown in current global sea surface turbulent flux datasets. This study uses synergized data of satellite global radiation, precipitation, and other meteorological variable observations to estimate sea surface turbulent heat fluxes. The radiation observations are made by the satellite Clouds and the Earth’s Radiant Energy System (CERES) sensors, while the global precipitation data is from NASA’s satellite Global Precipitation Climatology Project (GPCP). Other data includes satellite sea surface water temperature and wind speed observations. These datasets are obtained from a wide range of space sensors from passive to active instruments and from visible and near infrared to thermal infrared and microwave spectral sounders. One significant feature of these datasets is that they have multi-decades long climate records. Top-of-atmosphere (TOA) radiation and its anomaly represents the net heat energy input to the climate system, and the oceanic precipitation and ocean-land moisture transport can be used to quantify sea surface latent heat release. Based on the synergized datasets and the principle of global water and energy balances, global mean turbulent heat fluxes are estimated. The turbulent heat anomalies for the first two decades of the 21st century are, then, obtained mainly from global CERES radiation and GPCP precipitation anomalies, along with CERES derived ocean-land heat transports. Analysis indicates that the uncertainties in the estimated turbulent flux anomalies may be reduced considerably. The results suggest a strong needs in synchronized and synergized observations of atmospheric radiation, precipitation, and oceanic meteorological variables for long-term climate studies.

Bing Lin↗

Dynamic Server-Based KML Code Generator Method for Level-of-Detail Traversal of Geospatial Data

Web-based geospatial client applications such as Google Earth and NASA World Wind must listen to data requests, access appropriate stored data, and compile a data response to the requesting client application. This process occurs repeatedly to support multiple client requests and application instances. Newer Web-based geospatial clients also provide user-interactive functionality that is dependent on fast and efficient server responses. With massively large datasets, server-client interaction can become severely impeded because the server must determine the best way to assemble data to meet the client applications request. In client applications such as Google Earth, the user interactively wanders through the data using visually guided panning and zooming actions. With these actions, the client application is continually issuing data requests to the server without knowledge of the server s data structure or extraction/assembly paradigm. A method for efficiently controlling the networked access of a Web-based geospatial browser to server-based datasets in particular, massively sized datasets has been developed. The method specifically uses the Keyhole Markup Language (KML), an Open Geospatial Consortium (OGS) standard used by Google Earth and other KML-compliant geospatial client applications. The innovation is based on establishing a dynamic cascading KML strategy that is initiated by a KML launch file provided by a data server host to a Google Earth or similar KMLcompliant geospatial client application user. Upon execution, the launch KML code issues a request for image data covering an initial geographic region. The server responds with the requested data along with subsequent dynamically generated KML code that directs the client application to make follow-on requests for higher level of detail (LOD) imagery to replace the initial imagery as the user navigates into the dataset. The approach provides an efficient data traversal path and mechanism that can be flexibly established for any dataset regardless of size or other characteristics. The method yields significant improvements in userinteractive geospatial client and data server interaction and associated network bandwidth requirements. The innovation uses a C- or PHP-code-like grammar that provides a high degree of processing flexibility. A set of language lexer and parser elements is provided that offers a complete language grammar for writing and executing language directives. A script is wrapped and passed to the geospatial data server by a client application as a component of a standard KML-compliant statement. The approach provides an efficient means for a geospatial client application to request server preprocessing of data prior to client delivery. Data is structured in a quadtree format. As the user zooms into the dataset, geographic regions are subdivided into four child regions. Conversely, as the user zooms out, four child regions collapse into a single, lower-LOD region. The approach provides an efficient data traversal path and mechanism that can be flexibly established for any dataset regardless of size or other characteristics.

Baxes, Gregory↗

An A-Train Climatology of Extratropical Cyclone Clouds

Extratropical cyclones (ETCs) are the main purveyors of precipitation in the mid-latitudes, especially in winter, and have a significant radiative impact through the clouds they generate. However, general circulation models (GCMs) have trouble representing precipitation and clouds in ETCs, and this might partly explain why current GCMs disagree on to the evolution of these systems in a warming climate. Collectively, the A-train observations of MODIS, CloudSat, CALIPSO, AIRS and AMSR-E have given us a unique perspective on ETCs: over the past 10 years these observations have allowed us to construct a climatology of clouds and precipitation associated with these storms. This has proved very useful for model evaluation as well in studies aimed at improving understanding of moist processes in these dynamically active conditions. Using the A-train observational suite and an objective cyclone and front identification algorithm we have constructed cyclone centric datasets that consist of an observation-based characterization of clouds and precipitation in ETCs and their sensitivity to large scale environments. In this presentation, we will summarize the advances in our knowledge of the climatological properties of cloud and precipitation in ETCs acquired with this unique dataset. In particular, we will present what we have learned about southern ocean ETCs, for which the A-train observations have filled a gap in this data sparse region. In addition, CloudSat and CALIPSO have for the first time provided information on the vertical distribution of clouds in ETCs and across warm and cold fronts. We will also discuss how these observations have helped identify key areas for improvement in moist processes in recent GCMs. Recently, we have begun to explore the interaction between aerosol and cloud cover in ETCs using MODIS, CloudSat and CALIPSO. We will show how aerosols are climatologically distributed within northern hemisphere ETCs, and how this relates to cloud cover.

clouds↗

NanoPSD: A software for automatic detection of Nano-Particle Shape Distribution in electron microscopy images

Accurate quantification of the size and morphology of nanoparticles from electron microscopy (EM) images is essential to understand growth mechanisms, surface reactivity, and functional behavior in nanoscale materials. Manual analysis remains slow, subjective, and difficult to reproduce in large datasets. We introduce NanoPSD (Nano-Particle Shape Distribution), an open-source and fully automated framework for quantitative particle detection and morphology analysis from EM images. NanoPSD integrates adaptive contrast enhancement, polarity-agnostic scale-bar detection, Optical Character Recognition (OCR)-based calibration, and classical segmentation via Otsu thresholding with morphological refinement. Particle contours are used to extract geometric descriptors, including equivalent circular diameter, aspect ratio, circularity, and solidity, enabling automated classification into spherical, rod-like, and aggregate morphologies. The framework supports both single-image and batch processing, generating publication-quality visualizations, LaTeX-ready tables, and structured comma-separated values (CSV) datasets. As a demonstration, we applied NanoPSD to plasma-synthesized nanoparticle samples diagnosed via transmission electron microscopy (TEM). The code produced statistically robust size and morphology distributions spanning a few to tens of nanometers with minimal user supervision. The pipeline demonstrates high reproducibility and scalability, processing large image collections with consistent calibration and output formatting. Its modular design enables seamless integration of future deep-learning-based segmentation models, providing a pathway toward intelligent, data-driven electron microscopy analysis.

36 MATERIALS SCIENCE↗

Interoperable Map Services with Performance Tuning for Earth Science Data through API-Tiles and Dynamic API-Styles

NASA’s Goddard Earth Sciences Data and Information Services Center (GES DISC) provides access to a wide range of global climate data from various satellite missions and models. However, the visualization and analysis of these data can be challenging due to their large volume, complex structure, and diverse formats. This study presents the implementation of interoperable map services (API-Maps) with performance tuning using API-Tiles and dynamic API-Styles. API-Maps is a standard for defining and exposing map services through RESTful (representational state transfer) APIs (application programming interfaces). API-Tiles is a technique for generating and delivering map tiles on demand from any data source. API-Styles is a method for dynamically applying styles to map tiles based on user preferences or data attributes. The use of API-Tiles and dynamic API-Styles enhances the performance and scalability of the map services, allowing for smooth and interactive visualization of large datasets. Two types of Earth Science data sources from the NASA GES DISC are used in the experiment: regularly gridded data, such as Global Precipitation Measurement (GPM) precipitation data, and low processing level data, such as low-level data of atmospheric composite measurements from the TROPOspheric Monitoring Instrument (TROPOMI) mission. Re-gridding of swath data (low level data - e.g. Level 2) of atmospheric composites (e.g. TROPOMI products, such as nitrogen dioxide, ozone and aerosol optical depth) is applied to enable the Web-based, interoperable, tiled, and styled mapping (rendering) services of such data. The results demonstrate the effectiveness of the proposed approach in providing fast and efficient access to Earth science data through interoperable map services.

Geographic Information System↗

Simplifying NASA Earth Science Data and Information Access Through Natural Language Processing Based Data Analysis and Visualization

NASA Earth science data collected from satellites, model assimilation, airborne missions, and field campaigns, are large, complex and evolving. Such characteristics pose great challenges for end users (e.g., Earth science and applied science users, students, citizen scientists), particularly for those who are unfamiliar with NASA's EOSDIS and thus unable to access and utilize datasets effectively. For example, a novice user may simply ask: what is the total rainfall for a flooding event in my county yesterday? For an experienced user (e.g., algorithm developer), a question can be: how did my rainfall product perform, compared to ground observations, during a flooding event? Nonetheless, with rapid information technology development such as natural language processing, it is possible to develop simplified Web interfaces and back-end processing components to handle such questions and deliver answers in terms of text, data, or graphic results directly to users.In this presentation, we describe the main challenges for end users with different levels of expertise in accessing and utilizing NASA Earth science data. Surveys reveal that most non-professional users normally do not want to download and handle raw data as well as conduct heavy-duty data processing tasks. Often they just want some simple graphics or data for various purposes. To them, simple and intuitive user interfaces are sufficient because complicated ones can be difficult and time-consuming to learn. Professionals also want such interfaces to answer many questions from datasets. One solution is to develop a natural language based search box like Google and the search results can be text, data, graphics and more. Now the challenge is, with natural language processing, can we design a system to process a scientific question typed in by a user? In this presentation, we describe our plan for such a prototype. The workflow is: 1) extract needed information (e.g., variables, spatial and temporal information, processing methods, etc.) from the input, 2) process the data in the backend, and 3) deliver the results (data or graphics) to the user.

Liu, Zhong↗

A Step-by-Step Protocol from METASPACE to Biological Interpretation

Mass spectrometry imaging (MSI) represents an exceptional tool for exploring complex biological systems spatially at the molecular level. However, due to its multidimensional nature and large-scale data output, it presents considerable challenges when it comes to extracting meaningful biological insights. Recent advancements, such as the METASPACE platform, have enabled researchers to efficiently process, annotate, and interpret MSI datasets by leveraging machine learning and cloud-based infrastructure. In this tutorial, we present a detailed and user-friendly R-pipeline designed to help METASPACE users navigate untargeted metabolomic annotations and transform them into practical insights about their biological systems. By combining METASPACE annotations with rapid R-based screening, this workflow not only streamlined the analytical process but also enhanced the understanding of spatial molecular distribution, especially for complex systems. Here, this easy-to-follow approach has the potential for applications in diagnostics, drug discovery, environmental and ecological processes, and more. We envision this pipeline to be particularly useful for newcomers to the field of MSI and

Moreno Pedraza, Abigail↗

Tractometry of the Human Connectome Project: resources and insights

The Human Connectome Project (HCP) has become a keystone dataset in human neuroscience, with a plethora of important applications in advancing brain imaging methods and an understanding of the human brain. We focused on tractometry of HCP diffusion-weighted MRI (dMRI) data. We used an open-source software library (pyAFQ; https://yeatmanlab.github.io/pyAFQ) to perform probabilistic tractography and delineate the major white matter pathways in the HCP subjects that have a complete dMRI acquisition (n = 1,041). We used diffusion kurtosis imaging (DKI) to model white matter microstructure in each voxel of the white matter, and extracted tract profiles of DKI-derived tissue properties along the length of the tracts. We explored the empirical properties of the data: first, we assessed the heritability of DKI tissue properties using the known genetic linkage of the large number of twin pairs sampled in HCP. Second, we tested the ability of tractometry to serve as the basis for predictive models of individual characteristics (e.g., age, crystallized/fluid intelligence, reading ability, etc.), compared to local connectome features. To facilitate the exploration of the dataset we created a new web-based visualization tool and use this tool to visualize the data in the HCP tractometry dataset. Finally, we used the HCP dataset as a test-bed for a new technological innovation: the TRX file-format for representation of dMRI-based streamlines. We released the processing outputs and tract profiles as a publicly available data resource through the AWS Open Data program's Open Neurodata repository. We found heritability as high as 0.9 for DKI-based metrics in some brain pathways. We also found that tractometry extracts as much useful information about individual differences as the local connectome method. We released a new web-based visualization tool for tractometry—“Tractoscope” (https://nrdg.github.io/tractoscope). We found that the TRX files require considerably less disk space-a crucial attribute for large datasets like HCP. In addition, TRX incorporates a specification for grouping streamlines, further simplifying tractometry analysis.

59 BASIC BIOLOGICAL SCIENCES↗

GOLEM: GOld standard for Learning and Evaluation of Motifs

Motifs are distinctive, recurring, widely used idiom-like words or phrases, often originating from folklore, whose meaning is anchored in a narrative and have a significance as communicative devices across a wide range of media, including news, literature, and propaganda. Many motifs concisely imply a large constellation of culturally relevant information, and their broad usage suggests their cognitive importance as touchstones of cultural knowledge. As such, their detection is a step towards culturally aware natural language processing. We present GOLEM (GOld standard for Learning and Evaluation of Motifs) a dataset of English news articles, opinion pieces, and broadcast transcripts annotated for motific information. The dataset identifies 25,737 motif candidates across 34 motif types drawn from three cultural or national groups: Jewish, Irish, and Puerto Rican. The dataset contains 2,024,141 words split into 25,737 text snippets drawn from 8,073 articles. Each motif candidate is labeled according to a scheme which identifies the type of usage (motific, referential, eponymic, or unrelated), resulting in 1,743 actual motific instances in the data. Annotation was performed by individuals identifying as members of each group and achieved a Fleiss’ kappa (?) of > 0.55. In addition to the data, we demonstrate that classification of the candidate type is a challenging task for Large Language Models (LLMs) using a few-shot approach; recent models such as T5, FLAN-T5, GPT-2, and Llama 2 (7B) achieved a performance of 41% accuracy at best, where the majority class accuracy is 41% and the average chance accuracy is 27%. These data will support development of new models and approaches for detecting (and reasoning about) motific information in text.

motif, culture, natural language, artificial intel↗

A New Approach to using a Cloud-Resolving Model to Study the Interactions between Clouds, Precipitation and Aerosols

Numerical cloud models, which are based the non-hydrostatic equations of motion, have been extensively applied to cloud-scale and mesoscale processes during the past four decades. Because cloud-scale dynamics are treated explicitly, uncertainties stemming from convection that have to be parameterized in (hydrostatic) large-scale models are obviated, or at least mitigated, in cloud models. Global models will use the non-hydrostatic framework when their horizontal resolution becomes about 10 kilometers, the theoretical limit for the hydrostatic approximation. This juncture will be reached one to two decades from now. Over the past generation, voluminous datasets on atmospheric convection have been accumulated from radar, instrumented aircraft, satellites, and rawinsonde measurements in field campaigns, enabling the detailed evaluation of models. Improved numerical methods have resulted in more accurate and efficient dynamical cores in models. Improvements have been made in the parameterizations of microphysical processes, radiation, boundary-layer effects, and turbulence; however, microphysical parameterizations remain a major source of uncertainty in all classes of atmospheric models. In recent years, exponentially increasing computer power has extended cloud-resolving-model integrations from hours to months, the number of computational grid points from less than a thousand to close to ten million. Three-dimensional models are now more prevalent. Much attention is devoted to precipitating cloud systems where the crucial 1-kilometer scales are resolved in horizontal domains as large as 10,000 kilometers in two dimensions, and 1,000 x 1,000 square kilometers in three-dimensions. Cloud models now provide statistical information useful for developing more realistic physically-based parameterizations for climate models and numerical weather prediction models. A review of developments and applications of cloud models in the past, present and future will be presented in this talk. In particular, a new approach to using cloud-resolving models to study the interactions between clouds, precipitation and aerosols will be presented.

Tao, Wei-Kuo↗

A New Approach to Using a Cloud-resolving Model to Study the Interactions Between Clouds, Precipitation and Aerosols

Numerical cloud models, which are based the non-hydrostatic equations of motion, have been extensively applied to cloud-scale and mesoscale processes during the past four decades. Because cloud-scale dynamics are treated explicitly, uncertainties stemming from convection that have to be parameterized in (hydrostatic) large-scale models are obviated, or at least mitigated, in cloud models. Global models will use the non-hydrostatic framework when their horizontal resolution becomes about 10 km, the theoretical limit for the hydrostatic approximation. This juncture will be reached one to two decades from now. Over the past generation, voluminous datasets on atmospheric convection have been accumulated from radar, instrumented aircraft, satellites, and rawinsonde measurements in field campaigns, enabling the detailed evaluation of models. Improved numerical methods have resulted in more accurate and efficient dynamical cores in models. Improvements have been made in the parameterizations of microphysical processes, radiation, boundary-layer effects, and turbulence; however, microphysical parameterizations remain a major source of uncertainty in all classes of atmospheric models. In recent years, exponentially increasing computer power has extended cloud-resolving-model integrations from hours to months, the number of computational grid points from less than a thousand to close to ten million. Three-dimensional models are now more prevalent. Much attention is devoted to precipitating cloud systems where the crucial 1-km scales are resolved in horizontal domains as large as l0,OOO km in two-dimensions, and 1,OOO x 1,OOO km2 in three-dimensions. Cloud models now provide statistical information useful for developing more realistic physically-based parameterizations for climate models and numerical weather prediction models. A review of developments and applications of cloud models in the past, present and future will be presented in this talk. In particular, a new approach to using cloud-resolving models to study the interactions between clouds, precipitation and aerosols will be presented.

Tao, Wei-Kuo↗

A New Approach to using a Cloud-Resolving Model to Study the Interactions between Clouds, Precipitation and Aerosols

Numerical cloud models, which are based the non-hydrostatic equations of motion, have been extensively applied to cloud-scale and mesoscale processes during the past four decades. Because cloud-scale dynamics are treated explicitly, uncertainties stemming from convection that have to be parameterized in (hydrostatic) large-scale models are obviated, or at least mitigated, in cloud models. Global models will use the non-hydrostatic framework when their horizontal resolution becomes about 10 km, the theoretical limit for the hydrostatic approximation. This juncture will be reached one to two decades from now. Over the past generation, voluminous datasets on atmospheric convection have been accumulated from radar, instrumented aircraft, satellites, and rawinsonde measurements in field campaigns, enabling the detailed evaluation of models. Improved numerical methods have resulted in more accurate and efficient dynamical cores in models. Improvements have been made in the parameterizations of microphysical processes, radiation, boundary-layer effects, and turbulence; however, microphysical parameterizations remain a major source of uncertainty in all classes of atmospheric models. In recent years, exponentially increasing computer power has extended cloud-resolving-model integrations from hours to months, the number of computational grid points from less than a thousand to close to ten million. Three-dimensional models are now more prevalent. Much attention is devoted to precipitating cloud systems where the crucial 1-km scales are resolved in horizontal domains as large as 10,000 km in two-dimensions, and 1,000 x 1,000 square kilometers in three-dimensions. Cloud models now provide statistical information useful for developing more realistic physically-based parameterizations for climate models and numerical weather prediction models. A review of developments and applications of cloud models in the past, present and future will be presented in this talk. In particular, a new approach to using cloud-resolving models to study the interactions between clouds, precipitation and aerosols will be presented.

Tao, Wei-Kuo↗

Using Multi-scale Modeling System to Study the Interactions between Clouds, Precipitation, Aerosols, Radiation and Land Surface

Numerical cloud models, which are based the non-hydrostatic equations of motion, have been extensively applied to cloud-scale and mesoscale processes during the past four decades. Because cloud-scale dynamics are treated explicitly, uncertainties stemming from convection that have to be parameterized in (hydrostatic) large-scale models are obviated, or at least mitigated, in cloud models. Global models will use the non-hydrostatic framework when their horizontal resolution becomes about 10 kilometers, the theoretical limit for the hydrostatic approximation. This juncture will be reached one to two decades from now. Over the past generation, voluminous datasets on atmospheric convection have been accumulated from radar, instrumented aircraft, satellites, and rawinsonde measurements in field campaigns, enabling the detailed evaluation of models. Improved numerical methods have resulted in more accurate and efficient dynamical cores in models. Improvements have been made in the parameterizations of microphysical processes, radiation, boundary-layer effects, and turbulence; however, microphysical parameterizations remain a major source of uncertainty in all classes of atmospheric models. In recent years, exponentially increasing computer power has extended cloud-resolving-model integrations from hours to months, the number of computational grid points from less than a thousand to close to ten million. Three-dimensional models are now more prevalent. Much attention is devoted to precipitating cloud systems where the crucial 1-kilometer scales are resolved in horizontal domains as large as 10,000 kilometers in two-dimensions, and 1,000 x 1,000 square kilometers in three-dimensions. Cloud models now provide statistical information useful for developing more realistic physically based parameterizations for climate models and numerical weather prediction models. It is also expected that NWP and mesoscale model can be run in grid size similar to cloud resolving model through nesting technique. A review of developments, improvements and applications of cloud models (GCE and WRF) at Goddard will be presented in this talk. In particular, a new approach to using multi-scale modeling system to study the interactions between clouds, precipitation, aerosols and land will be presented.

Tao, Wei-Kuo↗