Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “spatio-temporal analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

E-transit-bench: simulation platform for analyzing electric public transit bus fleet operations

When electrified transit systems make grid aware choices, improved social welfare is achieved by reducing grid stress, reducing system loss, and minimizing power quality issues. Electrifying transit fleet has numerous challenges like non availability of buses during charging, varying charging costs and so on, that are related the electric grid behavior. However, transit systems do not have access to the information about the co-evolution of the grid's power flow and therefore cannot account for the power grid's needs in its day-to-day operation. In this paper we propose a framework of transportation-grid co-simulation, analyzing the spatio-temporal interaction between the transit operations with electric buses and the power distribution grid. Real-world data for a day's traffic from Chattanooga city's transit system is simulated in SUMO and integrated with a realistic distribution grid simulation (using GridLAB-D) to understand the grid impact due to transit electrification. Charging information is obtained from the transportation simulation to feed into grid simulation to assess the impact of charging. We also discuss the impact to the grid with higher degree of transit electrification that further necessitates such an integrated transportation-grid co-simulation to operate the integrated system optimally. Our future work includes extending the platform for optimizing the charging and trip assignment operations.

Sen, Rishav↗

Image Mapping and Visual Attention on the Sensory Ego-Sphere

The Sensory Ego-Sphere (SES) is a short-term memory for a robot in the form of an egocentric, tessellated, spherical, sensory-motor map of the robot s locale. Visual attention enables fast alignment of overlapping images without warping or position optimization, since an attentional point (AP) on the composite typically corresponds to one on each of the collocated regions in the images. Such alignment speeds analysis of the multiple images of the area. Compositing and attention were performed two ways and compared: (1) APs were computed directly on the composite and not on the full-resolution images until the time of retrieval; and (2) the attentional operator was applied to all incoming imagery. It was found that although the second method was slower, it produced consistent and, thereby, more useful APs. The SES is an integral part of a control system that will enable a robot to learn new behaviors based on its previous experiences, and that will enable it to recombine its known behaviors in such a way as to solve related, but novel, task problems with apparent creativity. The approach is to combine sensory-motor data association and dimensionality reduction to learn navigation and manipulation tasks as sequences of basic behaviors that can be implemented with a small set of closed-loop controllers. Over time, the aggregate of behaviors and their transition probabilities form a stochastic network. Then given a task, the robot finds a path in the network that leads from its current state to the goal. The SES provides a short-term memory for the cognitive functions of the robot, association of sensory and motor data via spatio-temporal coincidence, direction of the attention of the robot, navigation through spatial localization with respect to known or discovered landmarks, and structured data sharing between the robot and human team members, the individuals in multi-robot teams, or with a C3 center.

Fleming, Katherine Achim↗

Evaluating a Priori Ozone Profile Information Used in TEMPO (Tropospheric Emissions: Monitoring of Pollution) Tropospheric Ozone Retrievals

A primary objective for TOLNet is the evaluation and validation of space-based tropospheric O3 retrievals from future systems such as the Tropospheric Emissions: Monitoring of Pollution (TEMPO) satellite. This study is designed to evaluate the tropopause-based O3 climatology (TB-Clim) dataset which will be used as the a priori profile information in TEMPO O3 retrievals. This study also evaluates model simulated O3 profiles, which could potentially serve as a priori O3 profile information in TEMPO retrievals, from near-real-time (NRT) data assimilation model products (NASA Global Modeling and Assimilation Office (GMAO) Goddard Earth Observing System (GEOS-5) Forward Processing (FP) and Modern-Era Retrospective analysis for Research and Applications version 2 (MERRA2)) and full chemical transport model (CTM), GEOS-Chem, simulations. The TB-Clim dataset and model products are evaluated with surface (0-2 km) and tropospheric (0-10 km) TOLNet observations to demonstrate the accuracy of the suggested a priori dataset and information which could potentially be used in TEMPO O3 algorithms. This study also presents the impact of individual a priori profile sources on the accuracy of theoretical TEMPO O3 retrievals in the troposphere and at the surface. Preliminary results indicate that while the TB-Clim climatological dataset can replicate seasonally-averaged tropospheric O3 profiles observed by TOLNet, model-simulated profiles from a full CTM (GEOS-Chem is used as a proxy for CTM O3 predictions) resulted in more accurate tropospheric and surface-level O3 retrievals from TEMPO when compared to hourly (diurnal cycle evaluation) and daily-averaged (daily variability evaluation) TOLNet observations. Furthermore, it was determined that when large daily-averaged surface O3 mixing ratios are observed (65 ppb), which are important for air quality purposes, TEMPO retrieval values at the surface display higher correlations and less bias when applying CTM a priori profile information compared to all other data products. The primary reason for this is that CTM predictions better capture the spatio-temporal variability of the vertical profiles of observed tropospheric O3 compared to the TB-Clim dataset and other NRT data assimilation models evaluated during this study.

Pollution↗

Discovering a reaction–diffusion model for Alzheimer’s disease by combining PINNs with symbolic regression

Misfolded tau proteins play a critical role in the progression and pathology of Alzheimer's disease. Recent studies suggest that the spatio-temporal pattern of misfolded tau follows a reaction-diffusion type equation. However, the precise mathematical model and parameters that characterize the progression of misfolded protein across the brain remain incompletely understood. Here, we use deep learning and artificial intelligence to discover a mathematical model for the progression of Alzheimer's disease using longitudinal tau positron emission tomography from the Alzheimer's Disease Neuroimaging Initiative database. Specifically, we integrate physics informed neural networks (PINNs) and symbolic regression to discover a reaction-diffusion type partial differential equation for tau protein misfolding and spreading. First, we demonstrate the potential of our model and parameter discovery on synthetic data. Then, we apply our method to discover the best model and parameters to explain tau imaging data from 46 individuals who are likely to develop Alzheimer's disease and 30 healthy controls. Our symbolic regression discovers different misfolding models f(c) for two groups, with a faster misfolding for the Alzheimer's group, f(c) = 0.23c 3 – 1.34c 2 + 1.11c, than for the healthy control group, f(c) = –c 3 + 0.62c 2 + 0.39c. Our results suggest that PINNs, supplemented by symbolic regression, can discover a reaction-diffusion type model to explain misfolded tau protein concentrations in Alzheimer's disease. Furthermore, we expect our study to be the starting point for a more holistic analysis to provide image-based technologies for early diagnosis, and ideally early treatment of neurodegeneration in Alzheimer's disease and possibly other misfolding-protein based neurodegenerative disorders.

60 APPLIED LIFE SCIENCES↗

Aggregation Tool to Create Curated Data albums to Support Disaster Recovery and Response

Despite advances in science and technology of prediction and simulation of natural hazards, losses incurred due to natural disasters keep growing every year. Natural disasters cause more economic losses as compared to anthropogenic disasters. Economic losses due to natural hazards are estimated to be around $6-$10 billion dollars annually for the U.S. and this number keeps increasing every year. This increase has been attributed to population growth and migration to more hazard prone locations such as coasts. As this trend continues, in concert with shifts in weather patterns caused by climate change, it is anticipated that losses associated with natural disasters will keep growing substantially. One of challenges disaster response and recovery analysts face is to quickly find, access and utilize a vast variety of relevant geospatial data collected by different federal agencies such as DoD, NASA, NOAA, EPA, USGS etc. Some examples of these data sets include high spatio-temporal resolution multi/hyperspectral satellite imagery, model prediction outputs from weather models, latest radar scans, measurements from an array of sensor networks such as Integrated Ocean Observing System etc. More often analysts may be familiar with limited, but specific datasets and are often unaware of or unfamiliar with a large quantity of other useful resources. Finding airborne or satellite data useful to a natural disaster event often requires a time consuming search through web pages and data archives. Additional information related to damages, deaths, and injuries requires extensive online searches for news reports and official report summaries. An analyst must also sift through vast amounts of potentially useful digital information captured by the general public such as geo-tagged photos, videos and real time damage updates within twitter feeds. Collecting and aggregating these information fragments can provide useful information in assessing damage in real time and help direct recovery efforts. The search process for the analyst could be made much more efficient and productive if a tool could go beyond a typical search engine and provide not just links to web sites but actual links to specific data relevant to the natural disaster, parse unstructured reports for useful information nuggets, as well as gather other related reports, summaries, news stories, and images. This presentation will describe a semantic aggregation tool developed to address similar problem for Earth Science researchers. This tool provides automated curation, and creates "Data Albums" to support case studies. The generated "Data Albums" are compiled collections of information related to a specific science topic or event, containing links to relevant data files (granules) from different instruments; tools and services for visualization and analysis; information about the event contained in news reports, and images or videos to supplement research analysis. An ontology-based relevancy-ranking algorithm drives the curation of relevant data sets for a given event. This tool is now being used to generate a catalog of Hurricane Case Studies at Global Hydrology Resource Center (GHRC), one of NASA's Distribute Active Archive Centers. Another instance of the Data Albums tool is currently being created in collaboration with NASA/MSFC's SPoRT Center, which conducts research on unique NASA products and capabilities that can be transitioned to the operational community to solve forecast problems. This new instance focuses on severe weather to support SPoRT researchers in their model evaluation studies

Ramachandran, Rahul↗

The Cool Flames Experiment

A space-based experiment is currently under development to study diffusion-controlled, gas-phase, low temperature oxidation reactions, cool flames and auto-ignition in an unstirred, static reactor. At Earth's gravity (1g), natural convection due to self-heating during the course of slow reaction dominates diffusive transport and produces spatio-temporal variations in the thermal and thus species concentration profiles via the Arrhenius temperature dependence of the reaction rates. Natural convection is important in all terrestrial cool flame and auto-ignition studies, except for select low pressure, highly dilute (small temperature excess) studies in small vessels (i.e., small Rayleigh number). On Earth, natural convection occurs when the Rayleigh number (Ra) exceeds a critical value of approximately 600. Typical values of the Ra, associated with cool flames and auto-ignitions, range from 104-105 (or larger), a regime where both natural convection and conduction heat transport are important. When natural convection occurs, it alters the temperature, hydrodynamic, and species concentration fields, thus generating a multi-dimensional field that is extremely difficult, if not impossible, to be modeled analytically. This point has been emphasized recently by Kagan and co-workers who have shown that explosion limits can shift depending on the characteristic length scale associated with the natural convection. Moreover, natural convection in unstirred reactors is never "sufficiently strong to generate a spatially uniform temperature distribution throughout the reacting gas." Thus, an unstirred, nonisothermal reaction on Earth does not reduce to that generated in a mechanically, well-stirred system. Interestingly, however, thermal ignition theories and thermokinetic models neglect natural convection and assume a heat transfer correlation of the form: q=h(S/V)(T(bar) - Tw) where q is the heat loss per unit volume, h is the heat transfer coefficient, S/V is the surface to volume ratio, and (T(bar) - Tw ) is the spatially averaged temperature excess. This Newtonian form has been validated in spatially-uniform, well-stirred reactors, provided the effective heat transfer coefficient associated with the unsteady process is properly evaluated. Unfortunately, it is not a valid assumption for spatially-nonuniform temperature distributions induced by natural convection in unstirred reactors. "This is why the analysis of such a system is so difficult." Historically, the complexities associated with natural convection were perhaps recognized as early as 1938 when thermal ignition theory was first developed. In the 1955 text "Diffusion and Heat Exchange in Chemical Kinetics", Frank-Kamenetskii recognized that "the purely conductive theory can be applied at sufficiently low pressure and small dimensions of the vessel when the influence of natural convection can be disregarded." This was reiterated by Tyler in 1966 and further emphasized by Barnard and Harwood in 1974. Specifically, they state: "It is generally assumed that heat losses are purely conductive. While this may be valid for certain low pressure slow combustion regimes, it is unlikely to be true for the cool flame and ignition regimes." While this statement is true for terrestrial experiments, the purely conductive heat transport assumption is valid at microgravity (mu-g). Specifically, buoyant complexities are suppressed at mu-g and the reaction-diffusion structure associated with low temperature oxidation reactions, cool flames and auto-ignitions can be studied. Without natural convection, the system is simpler, does not require determination of the effective heat transfer coefficient, and is a testbed for analytic and numerical models that assume pure diffusive transport. In addition, mu-g experiments will provide baseline data that will improve our understanding of the effects of natural convection on Earth.

Pearlman, Howard↗

Statistical framework to assess long-term spatio-temporal climate changes: East River mountainous watershed case study

Abstract Evaluation of long-term temporal and spatial climatic change in mountainous regions is a critical challenge because of the interactive effects of multiple land and climatic factors and processes. Here we present the application of the statistical framework to the assessment of changes of climatic conditions, using data from 17 meteorological stations across the East River watershed near Crested Butte, Colorado, USA, and spanning the period from 1966 to 2021. The framework is developed based on (1) a time-series analysis of daily, monthly, and yearly averaged meteorological parameters (temperature, relative humidity, precipitation, wind speed, etc.), (2) evaluation and time series analysis of potential evapotranspiration (ET o ), actual evapotranspiration (ET), aridity index (AI), standard precipitation index (SPI) and standard precipitation-evapotranspiration index (SPEI), and (3) a temporal-spatial climatic zonation of the studied area based on the hierarchical clustering and PCA analysis of the SPEI, because the SPEI can be considered an integrative characteristic of the changes of climatic conditions. The Budyko model, with the application of the Penman–Monteith equation for the estimation of ET o , was used to determine the ET. The time series analysis of the AI is used to identify the periods with energy limited and water limited conditions. Hierarchical clustering of site locations for the three temporal segments of the SPEI showed a significant temporal-spatial shifts, indicating that dynamic climatic processes drive zonation patterns. Therefore, the watershed climatic zonation requires periodic re-evaluation based on the structural time series analysis of meteorological and water balance data.

54 ENVIRONMENTAL SCIENCES↗

Insights from application of a hierarchical spatio-temporal model to an intensive urban black carbon monitoring dataset

Existing regulatory pollutant monitoring networks rely on a small number of centrally located measurement sites that are purposefully sited away from major emission sources. While informative of general air quality trends regionally, these networks often do not fully capture the local variability of air pollution exposure within a community. Recent technological advancements have reduced the cost of sensors, allowing air quality monitoring campaigns with high spatial resolution. The 100×100 black carbon (BC) monitoring network deployed 100 low-cost BC sensors across the 15 km 2 West Oakland, CA community for 100 days in the summer of 2017, producing a nearly continuous site-specific time series of BC concentrations which we aggregated to one-hour averages. Leveraging this dataset, we employed a hierarchical spatio-temporal model to accurately predict local spatio-temporal concentration patterns throughout West Oakland, at locations without monitors (average cross-validated hourly temporal R 2 =0.60). Using our model, we identified spatially varying temporal pollution patterns associated with small-scale geographic features and proximity to local sources. In a sub-sampling analysis, here we demonstrated that fine scale predictions of nearly comparable accuracy can be obtained with our modeling approach by using ~30% of the 100×100 BC network supplemented by a shorter-term high-density campaign.

54 ENVIRONMENTAL SCIENCES↗

Entropy-based feature selection for capturing impacts in Earth system models with abrupt forcing

This paper presents the development of a new entropy-based feature selection method for identifying and quantifying impacts. Here, impacts are defined as statistically significant differences in spatio-temporal fields when comparing datasets with and without an external forcing in an Earth system model. Temporal feature selection is performed by first computing the cross-fuzzy entropy to quantify similarity of patterns between two datasets and then applying changepoint detection to identify regions of statistically constant entropy. The method is used to capture temperate north surface cooling from a 9-member simulation ensemble of the Mt. Pinatubo volcanic eruption, which injected 10 Tg of SO 2 into the stratosphere. The results estimate a mean difference decrease in near surface air temperature of -0.560 K with a 99% confidence interval between -0.864 K and -0.257 K between April and November of 1992, one year following the eruption. A sensitivity analysis with decreasing SO 2 injection revealed that the impact is statistically significant at 5 Tg but not at 3 Tg. Using identified features, a dependency graph model based on a 9-day lag had significantly fewer nodes than a graph based on monthly means. Furthermore, this demonstrates our method’s ability to perform dimension reduction while still uncovering source-to-impact pathways.

Changepoint detection↗

Multiresolution convolutional autoencoders

Herein we propose a multi-resolution convolutional autoencoder (MrCAE) architecture that integrates and leverages three highly successful mathematical architectures: (i) multigrid methods, (ii) convolutional autoencoders and (iii) transfer learning. The method provides an adaptive, hierarchical architecture that capitalizes on a progressive training approach for multiscale spatio-temporal data. This framework allows for inputs across multiple scales: starting from a compact (small number of weights) network architecture and low-resolution data, our network progressively deepens and widens itself in a principled manner to encode new information in the higher resolution data based on its current performance of reconstruction. Basic transfer learning techniques are applied to ensure information learned from previous training steps can be rapidly transferred to the larger network. As a result, the network can dynamically capture different scaled features at different depths of the network. The performance gains of this adaptive multiscale architecture are illustrated through a sequence of numerical experiments on synthetic examples and real-world spatial-temporal data.

97 MATHEMATICS AND COMPUTING↗

Characterizing vertical upper ocean temperature structures in the European Arctic through unsupervised machine learning

In-situ observations of subsurface ocean temperatures are, in many regions, inconsistently distributed in time and space. These spatio-temporal inconsistencies in the observational network lead to difficulties in utilizing those observations effectively for ocean model evaluation or understanding larger-scale ocean characteristics. Model accuracy of subsurface ocean characteristics is especially important within regions that contain complex ocean structures. One such region is the European Arctic which not only contains several types of water masses with unique characteristics, but also wintertime sea ice coverage and complex bathymetry. This study presents an unsupervised neural networking technique that can be used in combination with traditional ocean model evaluation techniques to provide additional information on the accuracy of modeled vertical ocean temperature profiles. Self-organizing maps is an unsupervised machine learning technique that we apply to approximately twenty thousand Argo and CTD temperature profiles from 2012 to 2020 in the European Arctic to categorize the observed vertical ocean temperature structures in the top 150 m. The observed ocean profile categories, or neurons, defined by the self-organizing map show strong spatial and temporal dependencies. We then use the neuron weights, or the learned temperature profile structure of each neuron, to validate the spatial and temporal variability of modeled vertical temperature structures. This analysis gives us new insights about the model’s capabilities to reproduce specific vertical structures of the top-most ocean layer within different regions and seasons. Mapping modeled ocean temperature profiles onto the neuron-space of the observationally-defined self organized map highlights the potential of this method to advance our understanding of model deficiencies in that region.

54 ENVIRONMENTAL SCIENCES↗

Small-Magnitude Seismic Swarms in Central Utah (US): Interactions of Regional Tectonics, Local Structures and Hydrothermal Systems

Swarms in Central Utah are situated in the complex transition between the Basin and Range (BR) province and the Colorado Plateau. Transecting transverse structures, volcanic deposits, and hydrothermal systems complicate the extensional BR horst and graben structures and provide a multitude of plausible triggering mechanisms. Revisiting the catalog of the University of Utah Seismograph Stations (1981–2022), we analyze spatio-temporal patterns and characteristic features of seismic sequences. Swarms with alternating seismicity rates, bursts, and longer swarms with persistent moment release exhibit a remarkable diversity in temporal evolution. Swarm durations do not scale with cumulative seismic moment: swarms lasting less than 1 day can have similar cumulative seismic moments as month-long swarms. Here, we observe stationary swarms re-occurring for years (e.g., Mineral Mountains), as well as singular swarms in low-seismicity areas (e.g., activating a local structure). The swarms show a pronounced heterogeneity in triggering and driving mechanisms, observed in the detailed analysis of exemplary sequences (detections, relocations, moment tensors, waveform-based clustering, and repeater analysis). The 2022 Sevier Valley sequence activated a BR-related normal fault, the first resolved fault plane in the valley since 1983. The 2011 Circleville sequence is interpreted as a swarm triggered by mainshock-aftershock activity characterized by increasing magnitudes, changing rupture mechanisms, and a concentration of highly similar events in the second part of the sequence. By jointly discussing exemplary sequences and catalog statistics, we draw a comprehensive picture of swarm activity and its relation to geothermal and tectonic activity.

58 GEOSCIENCES↗

DiffESM: Conditional Emulation of Temperature and Precipitation in Earth System Models With 3D Diffusion Models

Earth system models (ESMs) are essential for understanding the interaction between human activities and the Earth's climate. However, the computational demands of ESMs often limit the number of simulations that can be run, hindering the robust analysis of risks associated with extreme weather events. While low-cost climate emulators have emerged as an alternative to emulate ESMs and enable rapid analysis of future climate, many of these emulators only provide output on at most a monthly frequency. This temporal resolution is insufficient for analyzing events that require daily characterization, such as heat waves or heavy precipitation. We propose using diffusion models, a class of generative deep learning models, to effectively downscale ESM output from a monthly to a daily frequency. Trained on a handful of ESM realizations, reflecting a wide range of radiative forcings, our DiffESM model takes monthly mean precipitation or temperature as input, and is capable of producing daily values with statistical characteristics close to ESM output. Combined with a low-cost emulator providing monthly means, this approach requires only a small fraction of the computational resources needed to run a large ensemble. We evaluate model behavior using a number of extreme metrics, showing that DiffESM closely matches the spatio-temporal behavior of the ESM output it emulates in terms of the frequency and spatial characteristics of phenomena such as heat waves, dry spells, or rainfall intensity.

54 ENVIRONMENTAL SCIENCES↗

Unsupervised Clustering of Microseismic Events and Focal Mechanism Analysis at the CO 2 Injection Site in Decatur, Illinois

Characterization of induced microseismicity at a carbon dioxide (CO 2 ) storage site is critical for preserving reservoir integrity and mitigating seismic hazards. We apply a multilevel machine learning (ML) approach that combines the nonnegative matrix factorization and hidden Markov model to extract spectral representations of microseismic events and cluster them to identify seismic patterns at the Illinois Basin-Decatur Project. Unlike traditional waveform correlation methods, this approach leverages spectral characteristics of first arrivals to improve event classification and detect previously undetected planes of weakness. By integrating ML-based clustering with focal mechanism analysis, we resolve small-scale fault structures that are below the detection limits of conventional seismic imaging. Our findings reveal temporal bursts of microseismicity associated with brittle failure, providing insights into the spatio-temporal evolution of fault reactivation during CO 2 injection. This approach enhances seismic monitoring capabilities at CO 2 injection sites by improving fault characterization beyond the resolution of standard geophysical surveys.

Willis, Rachel Marie [Sandia National Laboratories↗

SMART Deliverable 6.1.2a: Application of the ORION tool to the IBDP Carbon Storage Site

Forecasting and managing potential induced seismic activity is one of the challenges facing commercialscale geologic carbon sequestration (GCS), as well as other geologic energy extraction and byproduct disposal technologies. Historically, the process to develop robust, science-based forecasts of induced seismicity has required an integrated effort from experts in seismology, geomechanics, and reservoir engineering to manage data, develop and evaluate models of subsurface processes, and to calibrate and interpret the results from a range of models to understand site behavior relative to prescribed standards and in the context of uncertainty in geologic characterization data, forecasting models, and operational scenario uncertainty. The Operational Forecasting of Induced Seismicity (ORION) toolkit is an open-source, observation-based forecasting toolkit that is being co-developed by two U.S. DOE-funded initiatives: the National Risk Assessment Partnership (NRAP) and the Scienceinformed Machine Learning for Accelerating Real Time Decisions in Subsurface Applications (SMART) Initiative. ORION is designed to provide functionality to support decision making about seismic hazard analysis and risk management for GCS stakeholders ranging from the public to site operators to expert seismologists. The tool, which is written as open-source code in the Python programming language, is composed of a desktop graphical user interface (GUI) and an underlying forecasting engine. The forecasting engine uses available reservoir properties, well and fluid injection scenario details, and observed seismic catalog data as inputs to produce a set of temporal and spatio-temporal seismic forecasts.

58 GEOSCIENCES↗

Refractivity Observations from Radar Phase Measurements: The 22 May 2002 Dryline Case during IHOP Project

The dryline, often associated with the development of severe storms in the Southern Great Plains of the United States of America, is a boundary layer phenomenon that occurs when a warm and moist air mass from the Gulf of Mexico meets a hot and dry air mass from the southwest desert area. An accurate knowledge of the water vapor spatio-temporal variability in the lower part of the atmosphere is crucial for a better understanding of the evolution of the dryline. The tropospheric refractivity, directly related to water vapor content, is a proxy for the water vapor content of the troposphere. It has already been demonstrated that the refractivity and the refractivity vertical gradient can be jointly estimated from radar phase measurements. In fact, it has been shown that using kriging interpolation techniques, accurate refractivity maps within the coverage area of the radar can be obtained with high temporal resolution. In this paper, a detailed analysis of the time series of radar-based refractivity maps obtained during a dryline that occurred on the afternoon of 22 May 2002 during the International H 2 O Project is presented. Comparisons between the time series of radar refractivity maps, obtained with the NCAR S-Pol radar, and the refractivity measurements derived from automatic ground-based weather stations and the AERI instrument, placed at different locations within the coverage area of the NCAR S-Pol radar, demonstrate the accuracy of radar refractivity estimates even for highly variable conditions, both in time and space, in the troposphere. Correlation coefficients higher than 0.95 are obtained in all weather station locations. Regarding the RMSE, errors less than 6 N-units are obtained for all cases, being even as low as 2.92 N-units at some locations.

54 ENVIRONMENTAL SCIENCES↗

Optimizing tertiary storage organization and access for spatio-temporal datasets

We address in this paper data management techniques for efficiently retrieving requested subsets of large datasets stored on mass storage devices. This problem represents a major bottleneck that can negate the benefits of fast networks, because the time to access a subset from a large dataset stored on a mass storage system is much greater that the time to transmit that subset over a network. This paper focuses on very large spatial and temporal datasets generated by simulation programs in the area of climate modeling, but the techniques developed can be applied to other applications that deal with large multidimensional datasets. The main requirement we have addressed is the efficient access of subsets of information contained within much larger datasets, for the purpose of analysis and interactive visualization. We have developed data partitioning techniques that partition datasets into 'clusters' based on analysis of data access patterns and storage device characteristics. The goal is to minimize the number of clusters read from mass storage systems when subsets are requested. We emphasize in this paper proposed enhancements to current storage server protocols to permit control over physical placement of data on storage devices. We also discuss in some detail the aspects of the interface between the application programs and the mass storage system, as well as a workbench to help scientists to design the best reorganization of a dataset for anticipated access patterns.

Chen, Ling Tony↗

Limiting Data Friction by Reducing Data Download Using Spatiotemporally Aligned Data Organization Through STARE

Current data processing practice limits the volume and variety of relevant geoscience data that can practically be applied to important problems. File archives in centralized data centers are the principal means by which Earth Science data are accessed. This approach, however, requires laborious search, retrieval, and eventual customization/adaptation for the data to be used. Such fractionation makes it even more difficult to share outcomes, i.e. research artifacts and data products, hampering reusability and repeatability, since end users generally have their own research agenda and preferences as well as scarce resources. Thus, while finding and downloading data files from central data centers are already costly for end users working in their own field, using data products from other disciplines rapidly becomes prohibitive. This curtails scientific productivity, limits avenues of study, and endangers quality and reproducibility. The Spatio-Temporal Adaptive Resolution Encoding (STARE) is a unifying scheme that facilitates the indexing, access, and fusion of diverse Earth Science data. STARE implements an innovative encoding of geo-spatiotemporal information, originally developed for aligning datasets with diverse spatiotemporal characteristics in an array database. The spatial component of STARE recursively quadfurcates a root polyhedron, producing a hierarchical scheme for addressing geographic locations and regions. The temporal component of STARE uses conventional date-time units as an indexing hierarchy. The additional encoding of spatial and temporal resolution information in STARE enables comparisons and conditional selections across diverse datasets. Moreover, spatiotemporal set-operations, e.g. union and intersection, are mapped to efficient integer operations with STARE. Applied to existing data models (point, grid, spacecraft swath) and corresponding granules, STARE indexes provide a streamlined description usable as geo-spatiotemporal metadata. When coupled with large scale, distributed hardware and software, STARE-based data access reduces pre-analysis data preparation costs by offering a convenient means to align different datasets spatiotemporally without specialized effort in parallel computing or distributed data management.

Kuo, Kwo-Sen↗