Engineering PapersSearch

Engineering topics

Thomas, Brian A.

Publications and source records attributed to Thomas, Brian A..

Reusing Data and Metadata to Create New Metadata Through Machine-Learning & Other Programmatic Methods

Recent improvements in natural language processing (NLP) enable metadata to be created programmatically from reused original metadata or even the dataset itself. Transfer-learning applied to NLP has greatly improved performance and reduced training data requirements. In this talk, we’ll compare machine-generated metadata to human-generated metadata and discuss characteristics of metadata and data archives that affect suitability for machine-learning reuse of metadata. Where as human-generated metadata is often populated once, populated from the perspective of data supplier, populated by many individuals with different words for the same thing, and limited in length, machine-generated metadata can be updated any number of times, generated from the perspective of any user, constrained to a standardized set of terms that can be evolved over time, and be any length required. Machine-learning generated metadata offers benefits but also additional needs in terms of version control, process transparency, human-computer interaction, and IT requirements. As a successful example, we’ll discuss how a dataset of abstracts and associated human-tagged keywords from a standardized list of several thousand keywords were used to create a machine-learning model that predicted keyword metadata for open-source code projects on code.nasa.gov. We’ll also discuss a less successful example from data.nasa.gov to show how data archive architecture and characteristics of initial metadata can be strong controls on how easy it is to leverage programmatic methods to reuse metadata to create additional metadata.

Gosses, Justin

Advanced Astrophysics Discovery Technology in the Era of Data Driven Astronomy

Astrophysics is at the threshold of a new epoch in which increasinglycomplex, heterogeneous datasets will challenge our existing information infrastructure and traditional approaches to analysis. The rapid advancement of graphics processing units, compact field programmable gate arrays and dedicated artificial intelligence accelerator chips is now permitting the use of scientific methods, processes and algorithms to extract knowledge and insights from structured and unstructured data in ways never before seen. Miniaturization of spacecraft architectures and supporting infrastructure is opening new observing strategies and new discovery spaces for science. The community is just beginning to awaken to these imminent challenges as evidenced by their relative lack of emphasis in the New Worlds, New Horizons ASTRO2010 decadal survey, in the ExoPAG Science Analysis Group 11 report andin the formulation of the WFIRST Data Challenge. We suggest that the Astrophysics Science Division (ASD), which has clearly recognized this new epoch of rapidly evolving information technology, could be more affirmative in its approach. We offer a modest structural solution.

Barry, Richard K.