Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “scientific machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

A pattern recognition system for locating small volvanoes in Magellan SAR images of Venus

The Magellan data set constitutes an example of the large volumes of data that today's instruments can collect, providing more detail of Venus than was previously available from Pioneer Venus, Venera 15/16, or ground-based radar observations put together. However, data analysis technology has not kept pace with data collection and storage technology. Due to the sheer size of the data, complete and comprehensive scientific analysis of such large volumes of image data is no longer feasible without the use of computational aids. Our progress towards developing a pattern recognition system for aiding in the detection and cataloging of small-scale natural features in large collections of images is reported. Combining classical image processing, machine learning, and a graphical user interface, the detection of the 'small-shield' volcanoes (less than 15km in diameter) that constitute the most abundant visible geologic feature in the more that 30,000 synthetic aperture radar (SAR) images of the surface of Venus are initially targeted. Our eventual goal is to provide a general, trainable tool for locating small-scale features where scientists specify what to look for simply by providing examples and attributes of interest to measure. This contrasts with the traditional approach of developing problem specific programs for detecting Specific patterns. The approach and initial results in the specific context of locating small volcanoes is reported. It is estimated, based on extrapolating from previous studies and knowledge of the underlying geologic processes, that there should be on the order of 10(exp 5) to 10(exp 6) of these volcanoes visible in the Magellan data. Identifying and studying these volcanoes is fundamental to a proper understanding of the geologic evolution of Venus. However, locating and parameterizing them in a manual manner is forbiddingly time-consuming. Hence, the development of techniques to partially automate this task were undertaken. The primary constraints for this particular problem are that the method must be reasonably robust and fast. Unlike most geological features, the small volcanoes of Venus can be ascribed to a basic process that produces features with a short list of readily defined characteristics differing significantly from other surface features on Venus. For pattern recognition purposes the relevant criteria include (1) a circular planimetric outline, (2) known diameter frequency distribution from preliminary studies, (3) a limited number of basic morphological shapes, and (4) the common occurrence of a single, circular summit pit at the center of the edifice.

Burl, M. C.↗

Planning for rover opportunistic science

The Mars Exploration Rover Spirit recently set a record for the furthest distance traveled in a single sol on Mars. Future planetary exploration missions are expected to use even longer drives to position rovers in areas of high scientific interest. This increase provides the potential for a large rise in the number of new science collection opportunities as the rover traverses the Martian surface. In this paper, we describe the OASIS system, which provides autonomous capabilities for dynamically identifying and pursuing these science opportunities during longrange traverses. OASIS uses machine learning and planning and scheduling techniques to address this goal. Machine learning techniques are applied to analyze data as it is collected and quickly determine new science gods and priorities on these goals. Planning and scheduling techniques are used to alter the behavior of the rover so that new science measurements can be performed while still obeying resource and other mission constraints. We will introduce OASIS and describe how planning and scheduling algorithms support opportunistic science.

artificial intelligence↗

High School Citizen Scientists Use AI/ML to Predict Intra-Ocular Pressure From Gene Expression Data for Spaceflown Mice

Artificial Intelligence (AI) and Machine Learning (ML) have increasingly become pivotal in biological and biomedical research, largely due to the culture of open data sharing and its associated benefits. The methodologies inherent in AI/ML are particularly adept at identifying and forecasting biological phenotypes from the vast amounts of data generated by next-generation sequencing technologies. These techniques offer substantial promise for advancing research in space biosciences and for the development of automated systems for monitoring space health. Nevertheless, there are crucial aspects to consider when training, validating, and testing machine learning models in both biological research and clinical contexts. It is essential that Open Science principles, including data sharing and the availability of open-source code, are complemented by high-quality, publicly accessible training resources. These resources should focus on best practices and include modules based on real-world scientific cases and data to ensure that future AI/ML practitioners gain practical experience with genuine problems. Addressing this knowledge gap, we have designed, developed, and delivered both interactive and self-paced training programs for citizen scientists worldwide, enabling them to utilize AI/ML for space biology research. This initiative was made possible through generous funding from a Transformation to Open Science Training grant. The interactive training sessions, conducted this summer, utilized AI/ML techniques to analyze data from the Open Science Data Repository, specifically targeting the effects of spaceflight on ocular structure and function. The dataset OSD-583, from the Rodent Research 9 mission, provides experimental data detailing the ocular responses of mice subjected to a 35-day spaceflight, compared with ground control counterparts. Using OSD-583 as observational data, our summer training participants applied AI/ML methods to predict intraocular pressure from RNA-seq data and identify the genes most predictive of the observed responses. Further analysis through pathway enrichment and gene set enrichment revealed that these genes are involved in molecular and cellular processes contributing to retinal degeneration.

James Casaletto↗

SatNet: A Benchmark for Satellite Scheduling Optimization

Satellites provide essential services such as networking and weather tracking, and the number of near-earth and deep space satellites are expected to grow rapidly in the coming years. Communications with terrestrial ground stations is one of the critical functionalities of any space mission. Satellite scheduling is a problem that has been scientifically investigated since the 1970s. A central aspect of this problem is the need to consider resource contention and satellite visibility constraints as they require line of sight. Due to the combinatorial nature of the problem, prior solutions such as linear programs and evolutionary algorithms require extensive compute capabilities to output a feasible schedule for each scenario. Machine learning based scheduling can provide an alternative solution by training a model with historical data and generating a schedule quickly with model inference. We present SatNet, a benchmark for satellite scheduling optimization based on historical data from the NASA Deep Space Network. We propose formulation of the satellite scheduling problem as a Markov Decision Process and use reinforcement learning (RL) policies to generate schedules. The nature of constraints imposed by SatNet differ from other combinatorial optimization problems such as vehicle routing studied in prior literature. Our initial results indicate that RL is an alternative optimization approach that can generate candidate solutions of comparable quality to existing state-of-the-practice results. However, we also find that RL policies overfit to the training dataset and do not generalize well to new data, thereby necessitating continued research on reusable and generalizable agents.

Wilson, Brian↗

The Impact of Dimensionality Reduction of Ion Counts Distributions on Preserving Moments, With Applications to Data Compression

The field of space physics has a long history of utilizing dimensionality reduction methods to distill data, including but not limited to spherical harmonics, the Fourier Transform, and the wavelet transform. Here, we present a technique for performing dimensionality reduction on ion counts distributions from the Multiscale Mission/Fast Plasma Investigation (MMS/FPI) instrument using a data-adaptive method powered by neural networks. This has applications to both feeding low-dimensional parameterizations of the counts distributions into other machine learning algorithms, and the problem of data compression to reduce transmission volume for space missions. The algorithm presented here is lossy, and in this work, we present the technique of validating the reconstruction performance with calculated plasma moments under the argument that preserving the moments also preserves fluid-level physics, and in turn a degree of scientific validity. The method presented here is an improvement over other lossy compressions in loss-tolerant scenarios like the Multiscale Mission/Fast Plasma Investigation Fast Survey or in non-research space weather applications.

D. da Silva↗

Improving GES Disc Data Search and Discovery Through AI Metadata Augmentation

NASA’s Goddard Earth Science (GES) Data and Information Services Center (DISC) is one of twelve data centers in NASA's Science Mission Directorate (SMD), providing vital earth science data to a diverse user base. To enhance the discoverability of this data, GES DISC employs a keyword search system, which leverages scientific keywords embedded in dataset metadata. However, the evolving nature of scientific applications of our data necessitates regular review and augmentation of these keywords. To address this, we developed a service to automatically predict missing science keywords in the metadata. This service constructs a knowledge graph from the latest GES DISC metadata within NASA’s Common Metadata Repository (CMR). Using an open-source library, we trained a machine learning model to predict absent science keywords in the metadata. Our preliminary results indicate that the model has high levels of accuracy at predicting science keywords in the dataset metadata when exposed to data not included in its training. These predicted keywords were then evaluated by GES DISC data curation scientists and compared against other AI tools for metadata augmentation. We aim to enhance the overall usability and accessibility of NASA’s earth science data by implementing this tool in our data curation processes.

Kendall Gilbert↗

Statistical Classification of Biosignature Information using Multiple Instrument Observations

The accurate identification of biosignatures (indications of life) from data taken from remote or in situ planetary exploration is one of the most important challenges in astrobiology, the interdisciplinary field examining habitability and the potential for extraterrestrial life. This study employs machine learning algorithms to optimize the identification of biosignatures, with an emphasis on those which are agnostic to a specific biochemical basis. We exploit the wealth of terrestrial data available from biogenic and abiogenic systems to enhance efficient feature prioritization. Our dataset, pulled from public databases and laboratory recorded measurements, includes elemental abundance, isotopic fractionation, and VNIR/Raman spectra The data curation process included standardization for detection limits and ranges. Subsequent feature extraction yielded detailed inputs for machine learning, including combinations of elemental content, isotopic ratios, and parameters of spectral peaks and troughs. Feature significance was evaluated across diverse machine learning methodologies, such as k-nearest neighbors, logistic regression, Random Forest, support vector machines, and Gaussian Naïve Bayes, along with a combined voting classifier. We utilized Receiver Operating Characteristic Area Under the Curve (ROC AUC) across 2,000 50% test-train splits as a robust metric of model performance. Results revealed a promising ROC AUC of 0.853 for the combined voting classifier. Removing elemental abundance data notably reduced model accuracy (13% decrease in AUC), highlighting its critical role in biosignature detection. Several other individual data features exhibited significance within their respective data types, offering additional granularity. This research fortifies the relevance of machine learning to astrobiology, potentially enhancing life detection missions by allowing algorithmic prioritization of high-interest samples for further investigation. Future work will refine data standardization, expand the dataset to include more terrestrial systems, and incorporate convolutional neural networks for spectral feature extraction. The potential for public data sharing is also under exploration, reinforcing our commitment to collective scientific advancement.

Statistical↗

Automatic Generation of Algorithms for the Statistical Analysis of Planetary Nebulae Images

Analyzing data sets collected in experiments or by observations is a Core scientific activity. Typically, experimentd and observational data are &aught with uncertainty, and the analysis is based on a statistical model of the conjectured underlying processes, The large data volumes collected by modern instruments make computer support indispensible for this. Consequently, scientists spend significant amounts of their time with the development and refinement of the data analysis programs. AutoBayes [GF+02, FS03] is a fully automatic synthesis system for generating statistical data analysis programs. Externally, it looks like a compiler: it takes an abstract problem specification and translates it into executable code. Its input is a concise description of a data analysis problem in the form of a statistical model as shown in Figure 1; its output is optimized and fully documented C/C++ code which can be linked dynamically into the Matlab and Octave environments. Internally, however, it is quite different: AutoBayes derives a customized algorithm implementing the given model using a schema-based process, and then further refines and optimizes the algorithm into code. A schema is a parameterized code template with associated semantic constraints which define and restrict the template s applicability. The schema parameters are instantiated in a problem-specific way during synthesis as AutoBayes checks the constraints against the original model or, recursively, against emerging sub-problems. AutoBayes schema library contains problem decomposition operators (which are justified by theorems in a formal logic in the domain of Bayesian networks) as well as machine learning algorithms (e.g., EM, k-Means) and nu- meric optimization methods (e.g., Nelder-Mead simplex, conjugate gradient). AutoBayes augments this schema-based approach by symbolic computation to derive closed-form solutions whenever possible. This is a major advantage over other statistical data analysis systems which use numerical approximations even in cases where closed-form solutions exist. AutoBayes is implemented in Prolog and comprises approximately 75.000 lines of code. In this paper, we take one typical scientific data analysis problem-analyzing planetary nebulae images taken by the Hubble Space Telescope-and show how AutoBayes can be used to automate the implementation of the necessary anal- ysis programs. We initially follow the analysis described by Knuth and Hajian [KHO2] and use AutoBayes to derive code for the published models. We show the details of the code derivation process, including the symbolic computations and automatic integration of library procedures, and compare the results of the automatically generated and manually implemented code. We then go beyond the original analysis and use AutoBayes to derive code for a simple image segmentation procedure based on a mixture model which can be used to automate a manual preproceesing step. Finally, we combine the original approach with the simple segmentation which yields a more detailed analysis. This also demonstrates that AutoBayes makes it easy to combine different aspects of data analysis.

Fischer, Bernd↗

Common Scientific and Technological Interests Between Astrobiology and Space Biology

The disciplines of astrobiology (AB) and space biology (SB) clearly have common interests, however they have not been pursued jointly. SB and AB are inextricably linked, both intellectually and technologically. They can now be effectively linked operationally. Cross-cutting joint collaborations will enhance innovation and increase cost effectiveness. Session topics include joint science questions, technologies, instrumentation, and missions. Examples include life detection, overlapping planetary protection concerns, biofilms, radiation, hyper- and hypogravity, applications of artificial intelligence and machine learning, interoperable databases, facilities (i.e., spacecraft, lunar surface efforts, simulation chambers, analog sites, etc.), training opportunities, and other topics relevant to AB and SB joint ventures. We welcome contributions on this very broad topical area to facilitate cross-fertilization of these disciplines that are of great importance to NASA.

astrobiology↗

Highlights of NASA’s Orbital Debris Program Office In Situ and Laboratory Measurements

NASA’s Orbital Debris Program Office (ODPO) maintains various returned spacecraft materials, capabilities, and facilities used for in situ and laboratory measurements that directly support orbital debris environmental models. In situ measurements include the analysis of exposed and returned hardware surfaces. These surfaces serve as passive sensors for the small-sized micrometeoroid and orbital debris (MMOD) flux below the sensitivity of ground-based radar and optical sensors. Various instruments and techniques are used to determine the size and depth of selected impact features, and – if feasible – the composition of the projectile material. Analysis of the impactor residues enables the differentiation of MM and OD for debris below 1 mm to support modeling the OD environment. In addition, projectiles identified as OD can be further differentiated in low-, medium-, and high- density impactors based on chemical analyses. In addition to in situ measurements, the ODPO has also worked in collaboration with the U.S. Space Force Space Systems Command (formerly the U.S. Air Force Space and Missile Systems Center), the Aerospace Corporation, and the University of Florida on a laboratory-based hypervelocity impact test, DebriSat, conducted at the Air Force Arnold Engineering Development Complex in 2014. The resulting data from this impact test series are being analyzed to assess the fragments’ sizes/masses, materials/densities, shapes, and other parameters of interest. The DebriSat project provides the data needed to update NASA’s breakup models and size estimation models using the simulated orbital breakup of a modern, low Earth orbit spacecraft. Ultimately, over 200,000 fragments from this impact test will be stored at NASA Johnson Space Center (JSC) and further analyzed by the ODPO. This project will also use machine learning techniques to infer physical parameters of fragments embedded in the soft-catch foam used in the impact experiment. Applied to X-ray imagery of the foam panels, these techniques promise to minimize human-in-the-loop processes for fragment extraction and physical characterization. A brief overview of this project and data collected will be presented. Lastly, the ODPO provides various capabilities hosted at NASA JSC for optical inspections and measurements using a variety of techniques and scientific instrumentation to support both in situ and laboratory measurements. The ODPO’s Optical Measurement Center (OMC) is an advanced facility for photometric and spectroscopic laboratory measurements of targets, including fragments from the DebriSat project. The OMC simulates telescopic observations by using space-like illumination conditions and source-target-sensor orientation techniques. Additionally, the OMC is uniquely equipped to acquire pseudo-bidirectional reflectance distribution data for broadband photometric measurements, thus removing aspect angle dependencies that can affect target size estimates using the optical size estimation model. Narrow-band surface material characterization using spectroscopic instrumentation gives insight into how the albedo parameter – also important in the optical size estimation model – may vary depending on the state of the material. The OMC also performs simulations of photometric measurements using optical ray-tracing software to model the OMC optical throughput. In addition to the OMC, the ODPO houses a start-of-the-art Fragment Analysis Facility that uses multiple microscopic inspection instruments to support in situ measurements and material characterization. An overview of both facilities will be highlighted in this paper.

Orbital Debris↗

Highlights of NASA’s Orbital Debris Program Office In Situ and Laboratory Measurements

NASA’s Orbital Debris Program Office (ODPO) maintains various returned spacecraft materials, capabilities, and facilities used for in situ and laboratory measurements that directly support orbital debris environmental models. In situ measurements include the analysis of exposed and returned hardware surfaces. These surfaces serve as passive sensors for the small-sized micrometeoroid and orbital debris (MMOD) flux below the sensitivity of ground-based radar and optical sensors. Various instruments and techniques are used to determine the size and depth of selected impact features, and – if feasible – the composition of the projectile material. Analysis of the impactor residues enables the differentiation of MM and OD for debris below 1 mm to support modeling the OD environment. In addition, projectiles identified as OD can be further differentiated in low-, medium-, and high- density impactors based on chemical analyses. In addition to in situ measurements, the ODPO has also worked in collaboration with the U.S. Space Force Space Systems Command (formerly the U.S. Air Force Space and Missile Systems Center), the Aerospace Corporation, and the University of Florida on a laboratory-based hypervelocity impact test, DebriSat, conducted at the Air Force Arnold Engineering Development Complex in 2014. The resulting data from this impact test series are being analyzed to assess the fragments’ sizes/masses, materials/densities, shapes, and other parameters of interest. The DebriSat project provides the data needed to update NASA’s breakup models and size estimation models using the simulated orbital breakup of a modern, low Earth orbit spacecraft. Ultimately, over 200,000 fragments from this impact test will be stored at NASA Johnson Space Center (JSC) and further analyzed by the ODPO. This project will also use machine learning techniques to infer physical parameters of fragments embedded in the soft-catch foam used in the impact experiment. Applied to X-ray imagery of the foam panels, these techniques promise to minimize human-in-the-loop processes for fragment extraction and physical characterization. A brief overview of this project and data collected will be presented. Lastly, the ODPO provides various capabilities hosted at NASA JSC for optical inspections and measurements using a variety of techniques and scientific instrumentation to support both in situ and laboratory measurements. The ODPO’s Optical Measurement Center (OMC) is an advanced facility for photometric and spectroscopic laboratory measurements of targets, including fragments from the DebriSat project. The OMC simulates telescopic observations by using space-like illumination conditions and source-target-sensor orientation techniques. Additionally, the OMC is uniquely equipped to acquire pseudo-bidirectional reflectance distribution data for broadband photometric measurements, thus removing aspect angle dependencies that can affect target size estimates using the optical size estimation model. Narrow-band surface material characterization using spectroscopic instrumentation gives insight into how the albedo parameter – also important in the optical size estimation model – may vary depending on the state of the material. The OMC also performs simulations of photometric measurements using optical ray-tracing software to model the OMC optical throughput. In addition to the OMC, the ODPO houses a start-of-the-art Fragment Analysis Facility that uses multiple microscopic inspection instruments to support in situ measurements and material characterization. An overview of both facilities will be highlighted in this paper.

Orbital Debris↗

Report on Workshop on Artificial Intelligence in Strategic Planning and Science Prioritization

This report details the observations from a two-day virtual workshop, held May 12-13, 2020, focused on whether, and how, artificial intelligence (AI) could assist humans in strategic planning, specifically in science and technology prioritization. The participants identified several “key challenges” that AI might tackle in this area. To further understand the value of these key challenges the workshop then developed related test cases that would demonstrate specifically how AI/machine learning (ML) could provide assistance to humans. Approximately 40 subject matter experts (SMEs), with backgrounds in AI, strategic planning for science, and scientific data, were gathered for the conference. This report collates the details of the output of the workshop. The “best” test cases include (in no particular order):Use of AI to assist in selecting Decadal Survey priorities. * Use of AI to identify new, or previously unidentified, science topics for prioritization. * Using AI to better label and increase discoverability of scientific literature and proposals. * Use of AI to enhance current observation capabilities for scientific missions. * Using AI to mitigate biases in selection of proposal reviewers and membership of advisory committees. Examination of these test cases indicates that Natural Language Processing (NLP) is a common capability found in most of the ”best” (top-rated) test cases and is a valuable, multi-purpose tool which enables ML in this area.

strategic planning↗

Space Flown Rodent Liver RNA Sequencing Data for Machine Learning in Space Biology Research

High-throughput nucleic acid sequencing (DNA-seq, RNA-seq) has become widespread in biomedical research due to the growing availability and affordability of these assays. Data analysis has been accelerated in recent years by the adoption of artificial intelligence (AI) and machine learning (ML) techniques by biomedical researchers. In space biology research, RNAseq datasets from space-flown experimental samples are critical for characterizing the gene expression aberrations associated with exposure to spaceflight stressors. However, space biological experiments tend to be very low sample size, so identifying proper AI/ML algorithms for sequencing data analysis is an ongoing challenge since these algorithms typically require large sample size. The NASA Science Mission Directorate (SMD) has started the “Benchmark Initiative for AI/ML”, focused on creating datasets meant for three main applications: 1) scientific benchmarking, which finds the best algorithm for a specific problem; 2) application benchmarking, which measures algorithm performance against a set of parameters; and 3) system benchmarking, which evaluates performance of hardware and software architecture. These scientific benchmarks consist of an AI-ready dataset and a reference implementation on a specific scientific question. In this work, we focused on generating standardized datasets to allow the scientific community to benchmark AI/ML algorithms in the domain of space biology. We present here a standardized, AI-ready, publicly available benchmark dataset for space biology RNA-seq data as a collaboration between the NASA AI4LS (Artificial Intelligence for Life Sciences) working group. and NASA’s SMD. This dataset consists of space-flown and ground control mouse liver found in the NASA GeneLab omics database. However, to amplify the small sample number (n=112 samples) for ML purposes, we employ Gaussian noise and a generative adversarial network to extend this dataset to 6,000 synthetic samples, matching the original gene expression characteristics.

James Casaletto↗

Automating sky object classification in astronomical survey images

We describe the application of machine classification techniques to the development of an automated tool for the reduction of a large scientific data set. The 2nd Palomer Observatory Sky Survey is nearly completed. This survey provides comprehensive coverage of the northern celestial hemisphere in the form of photographic plates. The plates are being transformed into digitized images whose quality will probably not be surpassed in the next ten to twenty years. The images are expected to contain on the order of 10(exp 7) galaxies and 10(exp 8) stars. Astronomers wish to determine which of these sky objects belong to various classes of galaxies and stars. The size of this data set precludes manual analysis. Our approach is to develop a software system which integrates the functions of independently developed techniques for image processing and data classification. Digitized sky images are passed through image processing routines to identify sky objects and to extract a set of features for each object. These routines are used to help select a useful set of attributes for classifying sky objects. Then GID3* and O-BTree, two inductive learning techniques, learn classification decision trees from examples. These classifiers will be used to process the rest of the data. This paper gives an overview of the machine learning techniques used, describes the details of our specific application, and reports the initial encouraging results. The results indicate that our approach is well-suited to the problem. The primary benefits of the approach are increased data reduction throughput and consistency of classification. The classification rules which are the product of the inductive learning techniques will form an object, examinable basis for classifying sky objects. A final, not to be underestimated benefit is that astronomers will be freed from the tedium of an intensely visual task to pursue more challenging analysis and interpretation problems based on automatically cataloged data.

Fayyad, Usama M.↗

A New Generation of Intelligent Trainable Tools for Analyzing Large Scientific Image Databases

In a variety of scientific disciplines two-dimensional digital image data is now relied on as a basic component of routine scientific investigation. The proliferation of image acquisition hardware such as multi-spectral remote-sensing platforms, medical imaging sensors, and high-resolution cameras have led to the widespread use of image data in fields such as atmospheric studies, planetary geology, ecology, agriculture, glacielogy, forestry, astronomy, diagnostic medicine, to name but a few.

machine↗

Transforming Science Prioritization Processes Using Artificial Intelligence

Artificial Intelligence (AI) and Machine Learning (ML) have potential to augment significantly the current labor-intensive processes of science prioritization, specifically by the National Academies’ Decadal Survey on behalf of NASA and NSF. Here we summarize what we believe to be the first exploratory demonstration-of-concept results from an application of AI/ML to Survey science prioritization. Specifically, we applied Latent Dirichlet Allocation (LDA) and Natural Language Processing (NLP) to reveal trends in published astrophysics research that may indicate science priorities and which could be applied to strategic planning. For the purpose of the work that we summarize here, AI/ML is able to analyze – that is, to “understand,” in a manner of speaking – a vast amount of text to reveal complex relationships among research topics, including the growth or decline of science community activities in those topics over time. We trained ourselves and AI/ML algorithms by using ~400,000 abstracts in the period 1998 to 2010 to “forecast” the Academies’ Astro2010 recommendations and compare with the solicited white papers. Comparing our results with actual Astro2010 recommendations allowed us to identify candidate metrics that better predicted the actual results of the Survey. We found, for example, that Compound Annual Growth Rate (CAGR) of papers published in a topic area is a good proxy measure for importance of this topic area of research. With this training complete, we identified candidate astrophysics astrophysics science priorities for the 2021+ period using the research during 2007 - 2019 . We conclude that appropriate application of AI can potentially significantly reduce the current workload of the Decadal Survey processes and reveal otherwise unrecognized characteristics in the body of astronomical research. We emphasize throughout the exploratory nature of our work, encouraging colleagues to pursue promising results further. Our most critical governing assumption was that increased (or decreased) research activity can be used to identify scientific or technology topic areas worthy of increased (or decreased) future emphasis. We discuss advantages, limitations, and recognize the “black box” nature of our technique. We note ethics issues associated, for example, with using AI/ML to reveal “hidden” meanings and biases in published work. Furthermore, inevitable improvements in AI may soon enable widespread and welcome identification of and advocacy for science and technology priorities by disparate and diverse groups and organizations. Consequently, we continue to urge a near-term, in-depth evaluation of appropriate applications of AI, including implications and consequences, as well as support for multiple follow-on assessments, of which ours is only a beginning.

Artificial Intelligence↗

Automated classification of scientific publications linked to GES DISC datasets

The data collections archived and distributedby the GES DISC NASA data center arewidely utilized for various Earth Science studies.As these collections are created, many researchworks are published regarding the collections, algorithms,validations and applications. SinceGES DISC collects these publications and providestheir citations for the users, it is helpful tocategorize them based on how they relate to the datasetsthey are associated with. Specifically,whether the publication that is linked to GES DISCdataset is using it for applicational research,or if it describes the algorithm for dataset creation,or the validation of the dataset, or providesthe general overview of the data collection. Currently,this process requires simple manuallabelling, and as such, may be possible to solve viaautomation. To approach this problem, wedeveloped machine learning classifiers to predictthe category a publication belongs to. We usedmanually labeled publications as training data forsupervised machine learning algorithms:Random Forest and Naive Bayes. We achieved classificationaccuracy that is substantially betterthan the baseline accuracy, thus greatly improvingthe efficiency of the publication internalanalysis.

Rohan Dayal↗