Engineering Papers⌕ Search

Engineering topics

Rahul Ramachandran

Publications and source records attributed to Rahul Ramachandran.

At least 109 records · Page 6

Modern Scientific Data Governance Framework

Science has entered the era of Big Data with new challenges related to data governance, stewardship, and management. The existing data governance practices must catch up to ensure proper data management. Existing data governance policies and stewardship best practices tend to be disconnected from operational data management practices and enforcement and mainly exist in well-meaning documents or reports. These governance policies are, at best, partially implemented and rarely monitored or audited. In addition, existing governance policies keep adding additional data management steps that require a human, ‘a data steward’, in the loop, and the cost of data management can no longer scale proportionately with the current and future increased data volume and complexity. The goal for developing an updated data governance framework is to modernize scientific data governance to the reality of Big data and align it with the current technology trends such as cloud computing and AI. The goals of this framework are two folds. One is to ensure thoroughness that the governance adequately covers the entire data life cycle. Two, provide a practical approach that offers a consistent and repeatable process for different projects. Three core principles ground this framework. First, focus on just enough governance and prevent data governance from becoming a roadblock toward the scientific process. Remove any unnecessary processes and steps. Second, automate data management steps where possible. Actively remove steps that require ‘human in the loop’ within the management process to be efficient and scale with increasing data. Third, all the processes should continually be optimized using quantified metrics to streamline the monitoring and auditing workflows.

Rahul Ramachandran↗

Foundation AI Models for Science

Foundation Models (FM) are AI models that are designed to replace a task or an application specific model. These FM can be applied to many different downstream applications. These FM are trained using self supervised techniques and can be built on any type of sequence data. The use of self supervised learning removes the hurdle for developing a large labeled dataset for training. Most FM use transformer architecture utilizes the notion of self attention which allows the network to model the influence of distant data points to each other both in space and time. The FM models exhibit emergent properties that are induced from the data. FM can be an important tool for science. The scale of these models results in better performance for different downstream applications and these applications show better accuracy over models built from scratch. FM drastically reduces the cost of entry to build different downstream applications both in time and effort. FM for selected science datasets such as optical satellite data, can accelerate applications ranging from data quality monitoring, feature detection and prediction. FM can make it easier to infuse AI into scientific research by removing the training data bottleneck and increasing the use of science data.

Manil Maskey↗

Observing Supraglacial Lakes Using Deep Learning and PlanetScope Imagery

Supraglacial lakes (SGL)s result from melt water accumulation in topographic depressions on the surface of glaciers. SGLs primarily affect glacial dynamics through a positive feedback loop in which the albedo-lowering effect of SGLs can escalate surface melt leading to increases in lake extent and depth, amplifying the afore mentioned albedo-lowering effect. The implications of accelerated glacial melt include increased sea level rise and modifications to ocean primary productivity. SGLs are critical indicators of surface melt and its downstream impacts and should be monitored efficiently. In situ observations and measurements of SGLs are time consuming, cost-prohibitive and difficult to scale. Earth observation data and machine learning enable scalable monitoring of SGLs through pattern detection and quantification of lake evolution over time [1]. This work presents a model developed by training a convolutional neural network with imagery and labels from NASA Operation IceBridge and predicting SGLs in high temporal and spatial resolution PlanetScope imagery.

Supraglacial lake↗

Revolutionizing Earth Science with Generalized AI Models

Foundation Models (FM) are generalized Artificial Intelligence (AI) models that are designed to replace a task or an application-specific model and can be used for many downstream applications. These FM can be built on any sequence data and are trained utilizing self-supervised approaches. The obstacle of creating a sizable labeled dataset for training is removed by using self-supervised learning. Most FM employ transformer design that takes advantage of the idea of self-attention, allowing the network to represent the impact of distant data points on one another in space and time. The FM models show emergent qualities that are induced from the data. FM can become a valuable tool for Earth science researchers. Due to the size of these models, downstream applications built fine-tuning these FM perform better and exhibit greater accuracy than models created from scratch. FM significantly lowers the entry barrier in terms of both the time and effort required to develop various downstream applications. For some scientific datasets, such as optical remote sensing data, FM can speed up processes like classification, object detection and prediction. By eliminating the training data bottleneck and maximizing the usage of science data, FM can make it simpler to integrate AI into scientific research. Initial results for three different FMs will be presented.

Rahul Ramachandran↗

NASA's Science Discovery Engine: Enabling Interdisciplinary Open Science

NASA is currently implementing capabilities to enable open science access. The Science Discovery Engine (SDE) is a major endeavor in this effort. The SDE supports discovery and access to complex, heterogeneous science data and information across science topical areas.

Kaylin Mclendon Bugbee↗

AI Foundation Models for Science: An Open Collaborative Initiative

Foundation Models (FMs), AI models designed to replace task-specific models, are increasingly being recognized for their versatility across numerous downstream applications. These models, trained using self-supervised techniques on any type of sequence data, circumvent the need for large annotated datasets, a major bottleneck in traditional AI model development. FMs can be applied to downstream tasks using few-shot learning and fine-tuning, significantly reducing the need for large labeled training datasets and computational resources. However, the development of FMs requires substantial resources, including access to data and compute power, expertise in the latest models, and specialized scientific knowledge for systematic evaluation. It is challenging for a single group to possess all these capabilities. To address this, NASA IMPACT has initiated an open collaborative effort, leveraging partnerships with the private sector and other groups within and outside NASA, to jointly build FMs. The overarching goal is to develop a consistent and collaborative approach to building FMs for high-value science datasets. This initiative has fostered collaboration within NASA and with external partners, including IBM Research, Clark University, DOE’s ORNL, ESA, and USGS. The effort focuses on identifying key datasets with a wide range of downstream applications, pretraining and building FMs using modified transformer architectures, evaluating compute infrastructure needs, and sharing models, pretraining and fine-tuning code, and data with the community. Furthermore, it aims to train the Earth science community to fine-tune these models for various downstream applications. Our initial effort resulted in the creation of a 100 million parameter HLS Geospatial Model within six months, which was released on HuggingFace. We are now expanding our scope to include data from weather and climate models and investigating multimodal models. We invite those interested in participating in this effort to join us by sharing their use cases, expertise, or data.

Rahul Ramachandran↗

Harnessing Large Language Models for Scientific Endeavors

The rapid proliferation of Large Language Models (LLMs) such as GPT, Bard, and Llama has revolutionized various sectors, including the scientific community. These models, with their potential to automate and augment tasks, are increasingly being recognized as both a valuable asset and a potential challenge in the realm of scientific research and data management. However, the current LLMs, primarily trained on general corpora, exhibit a limited understanding of scientific concepts and terminologies due to the lack of scientific corpus in their training data. Recognizing this gap, several groups are now advocating for the development of LLMs specifically tailored for scientific applications. A notable initiative in this direction is the Large Language Model effort initiated by NASA's CSDO. This endeavor aims to align LLM efforts across NASA’s Science Mission Directorate, develop a science-specific corpus and validation test set for model training, and create an encoder-only model for various downstream tasks. Moreover, the initiative also plans to develop a decoder-only model to explore the potential benefits and risks associated with a generative LLM for science. Lastly, the project aims to create a science evaluation suite, encompassing various categories of downstream scientific tasks, to serve as a benchmark for assessing the value of any LLM for future use. This presentation will provide an overview and current status of this ongoing initiative, highlighting its potential to reshape the use of LLMs in the scientific domain.

Rahul Ramachandran↗