Engineering Papers⌕ Search

Engineering topics

Devarakonda, Ranjeet

Publications and source records attributed to Devarakonda, Ranjeet.

Semi-automated Design of Artificial Intelligence Earth Science Models

Prediction and observation of water cycles involve not only patterns isolated in space and time, but rather modeling complex spatio-temporal relationships across multiple sources of data and domains. For instance, Evapotranspiration (ET) and Leaf Area Indexes (LAI) are two critical components in DOE’s Energy Exascale Earth System Model (E3SM). Accurate assessments of ET and LAI are critical for understanding hydrological processes, deforestation, crop yield, and irrigation impacts. However, current ET estimates for global simulations are available at very coarse spatial resolution. They are usually derived from satellite data based on broad plant functional types (PFT), which fail to capture the fine-scale variations due to change in vegetation type across the globe. Within this context and in light of the data-model integration challenges highlighted in the EESSD Strategic Plan, the new era of AI model development for geosciences calls for data-driven methods that provide domain scientists with estimations of parameters such as PFT and LAI in an efficient, interpretable, and easy-to-operate manner.

Ambrozio Dias, Philipe↗

FAIR Interfaces for Geospatial Scientific Data Searches

Several factors must be considered in designing a highly accurate, reliable, scalable, and user-friendly geospatial data search interfaces. This paper examines four critical questions that ought to be considered during design phase: (1) Is the search interface or API that provides the search capability useable by both humans and machines? (2) Are the results consistent and reliable? (3) Is the output response format free to use, community-defined, and non-propriety? (4) Does the API clearly state the usage clauses? This paper discusses how certain data repositories at the US Department of Energy's Oak Ridge National Laboratory apply FAIR data principles to enable geospatial searches and address the above-mentioned questions.

Devarakonda, Ranjeet↗

A Guide to Using GitHub for Developing and Versioning Data Standards and Reporting Formats

Abstract Data standardization combined with descriptive metadata facilitate data reuse, which is the ultimate goal of the Findable, Accessible, Interoperable, and Reusable (FAIR) principles. Community data or metadata standards are increasingly created through an approach that emphasizes collaboration between various stakeholders. Such an approach requires platforms for collaboration on the development process that centers on sharing information and receiving feedback. Our objective in this study was to conduct a systematic review to identify data standards and reporting formats that use version control for developing data standards and to summarize common practices, particularly in earth and environmental sciences. Out of 108 data standards and reporting formats identified in our review, 32 used GitHub as the version control platform, and no other platforms were used. We found no universally accepted methodology for developing and publishing data standards. Many GitHub repositories did not use key features that could help developers to gather user feedback, or to create and revise standards that build on previous work. We provide guidance for community‐driven standard development and associated documentation on GitHub based on a systematic review of existing practices.

54 ENVIRONMENTAL SCIENCES↗

AGU/AMS Abstract Search and Display Software

The AGU/AMS Abstract Search and Display Software is a standalone web application which enables the searching, storing, and displaying of abstracts featured at the annual American Geophysical Union (AGU) and American Meteorological Society (AMS) meetings. This application is designed for those who wish to host a standalone web application and feature a select subset of posters and talks scheduled for the AGU/AMS meetings. Please read the entirety of this README.md file before attempting to download and use the application. There are three views available via the UI: Lookup - enables searching and submitting posters for displaying on the summary view Manual Submission - allows individual manual submission of posters given a poster ID Summary - displays all posters submitted by users from the lookup view

Darnell, Wade↗

Enabling modern data discovery for atmospheric measurements

The Atmospheric Radiation Measurement (ARM) user facility is a US Department of Energy Office of Science user facility that is managed and operated through a collaborative effort led by nine US Department of Energy national laboratories. The ARM Data Center, located at Oak Ridge National Laboratory, is responsible for the timely collection, processing, and delivery of data products to the scientific community. The ARM Data Center holds more than 11,000 data products, including metadata collected from field campaigns, instruments, value-added products, and principal investigator–contributed data. These data sets are checked for successful transfer (for most data, this transfer is carried out automatically via the network; however, some of the largest data sets and some of the most remote sites require manual shipping of hard disks) and both the data and metadata are processed to a standard format, which is an ARM-standardized structure, via the Network Common Data Form. The Network Common Data Form is a self-describing binary format with many compatible software tools. Once processed, the data are cataloged, stored in the ARM Data Archive, and made discoverable through association with an array of metadata-characterizing information, such as location and measurement classification. These metadata enable powerful search capabilities through the ARM Data Center Data Discovery interface. This paper discusses the workflow of how the new discovery system has been redesigned from user requirements and how the data are distributed to the scientific community.

54 ENVIRONMENTAL SCIENCES↗

Mercury Toolset for Spatiotemporal Metadata

Mercury (http://mercury.ornl.gov) is a set of tools for federated harvesting, searching, and retrieving metadata, particularly spatiotemporal metadata. Version 3.0 of the Mercury toolset provides orders of magnitude improvements in search speed, support for additional metadata formats, integration with Google Maps for spatial queries, facetted type search, support for RSS (Really Simple Syndication) delivery of search results, and enhanced customization to meet the needs of the multiple projects that use Mercury. It provides a single portal to very quickly search for data and information contained in disparate data management systems, each of which may use different metadata formats. Mercury harvests metadata and key data from contributing project servers distributed around the world and builds a centralized index. The search interfaces then allow the users to perform a variety of fielded, spatial, and temporal searches across these metadata sources. This centralized repository of metadata with distributed data sources provides extremely fast search results to the user, while allowing data providers to advertise the availability of their data and maintain complete control and ownership of that data. Mercury periodically (typically daily) harvests metadata sources through a collection of interfaces and re-indexes these metadata to provide extremely rapid search capabilities, even over collections with tens of millions of metadata records. A number of both graphical and application interfaces have been constructed within Mercury, to enable both human users and other computer programs to perform queries. Mercury was also designed to support multiple different projects, so that the particular fields that can be queried and used with search filters are easy to configure for each different project.

Wilson, Bruce E.↗