Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “research data management”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

The crustal dynamics intelligent user interface anthology

The National Space Science Data Center (NSSDC) has initiated an Intelligent Data Management (IDM) research effort which has, as one of its components, the development of an Intelligent User Interface (IUI). The intent of the IUI is to develop a friendly and intelligent user interface service based on expert systems and natural language processing technologies. The purpose of such a service is to support the large number of potential scientific and engineering users that have need of space and land-related research and technical data, but have little or no experience in query languages or understanding of the information content or architecture of the databases of interest. This document presents the design concepts, development approach and evaluation of the performance of a prototype IUI system for the Crustal Dynamics Project Database, which was developed using a microcomputer-based expert system tool (M. 1), the natural language query processor THEMIS, and the graphics software system GSS. The IUI design is based on a multiple view representation of a database from both the user and database perspective, with intelligent processes to translate between the views.

Short, Nicholas M., Jr.↗

FAIR for AI: An interdisciplinary and international community building perspective

A foundational set of findable, accessible, interoperable, and reusable (FAIR) principles were proposed in 2016 as prerequisites for proper data management and stewardship, with the goal of enabling the reusability of scholarly data. The principles were also meant to apply to other digital assets, at a high level, and over time, the FAIR guiding principles have been re-interpreted or extended to include the software, tools, algorithms, and workflows that produce data. FAIR principles are now being adapted in the context of AI models and datasets. Here, we present the perspectives, vision, and experiences of researchers from different countries, disciplines, and backgrounds who are leading the definition and adoption of FAIR principles in their communities of practice, and discuss outcomes that may result from pursuing and incentivizing FAIR AI research. The material for this report builds on the FAIR for AI Workshop held at Argonne National Laboratory on June 7, 2022.

97 MATHEMATICS AND COMPUTING↗

System of experts for intelligent data management (SEIDAM)

A proposal to conduct research and development on a system of expert systems for intelligent data management (SEIDAM) is being developed. CCRS has much expertise in developing systems for integrating geographic information with space and aircraft remote sensing data and in managing large archives of remotely sensed data. SEIDAM will be composed of expert systems grouped in three levels. At the lowest level, the expert systems will manage and integrate data from diverse sources, taking account of symbolic representation differences and varying accuracies. Existing software can be controlled by these expert systems, without rewriting existing software into an Artificial Intelligence (AI) language. At the second level, SEIDAM will take the interpreted data (symbolic and numerical) and combine these with data models. at the top level, SEIDAM will respond to user goals for predictive outcomes given existing data. The SEIDAM Project will address the research areas of expert systems, data management, storage and retrieval, and user access and interfaces.

Goodenough, David G.↗

The development of an intelligent user interface for NASA's scientific databases

The National Space Science Data Center (NSSDC) has initiated an Intelligent Data Management (IDM) research effort which has as one of its components, the development of an Intelligent User Interface (IUI). The intent of the IUI effort is to develop a friendly and intelligent user interface service that is based on expert systems and natural language processing technologies. This paper presents the design concepts, development approach and evaluation of performance of a prototype Intelligent User Interface Subsystem (IUIS) supporting an operational database.

Campbell, William J.↗

The development of an intelligent user interface for NASA's scientific databases

The National Space Science Data Center (NSSDC) has initiated an Intelligent Data Management (IDM) research effort which has as one of its components, the development of an Intelligent User Interface (IUI). The intent of the IUI effort is to develop a friendly and intelligent user interface service that is based on expert systems and natural language processing technologies. This paper presents the design concepts, development approach and evaluation of performance of a prototype Intelligent User Interface Subsystem (IUIS) supporting an operational database.

Campbell, William J.↗

Harnessing Large Language Models for Scientific Endeavors

The rapid proliferation of Large Language Models (LLMs) such as GPT, Bard, and Llama has revolutionized various sectors, including the scientific community. These models, with their potential to automate and augment tasks, are increasingly being recognized as both a valuable asset and a potential challenge in the realm of scientific research and data management. However, the current LLMs, primarily trained on general corpora, exhibit a limited understanding of scientific concepts and terminologies due to the lack of scientific corpus in their training data. Recognizing this gap, several groups are now advocating for the development of LLMs specifically tailored for scientific applications. A notable initiative in this direction is the Large Language Model effort initiated by NASA's CSDO. This endeavor aims to align LLM efforts across NASA’s Science Mission Directorate, develop a science-specific corpus and validation test set for model training, and create an encoder-only model for various downstream tasks. Moreover, the initiative also plans to develop a decoder-only model to explore the potential benefits and risks associated with a generative LLM for science. Lastly, the project aims to create a science evaluation suite, encompassing various categories of downstream scientific tasks, to serve as a benchmark for assessing the value of any LLM for future use. This presentation will provide an overview and current status of this ongoing initiative, highlighting its potential to reshape the use of LLMs in the scientific domain.

Rahul Ramachandran↗

Earth science data study

The research proposed in this contract concerning investigations of existing and planned Earth Science and Applications Division (ESAD) data management systems and research into utilities for the access and display of scientific data products was completed. A summary of this work is provided.

Graves, Sara J.↗

Digital processing of mesoscale analysis and space sensor data

The mesoscale analysis and space sensor (MASS) data management and analysis system on the research computer system is presented. The MASS data base management and analysis system was implemented on the research computer system which provides a wide range of capabilities for processing and displaying large volumes of conventional and satellite derived meteorological data. The research computer system consists of three primary computers (HP-1000F, Harris/6, and Perkin-Elmer 3250), each of which performs a specific function according to its unique capabilities. The overall tasks performed concerning the software, data base management and display capabilities of the research computer system in terms of providing a very effective interactive research tool for the digital processing of mesoscale analysis and space sensor data is described.

Hickey, J. S.↗

System of Experts for Intelligent Data Management (SEIDAM)

It is proposed to conduct research and development on a system of expert systems for intelligent data management (SEIDAM). CCRS has much expertise in developing systems for integrating geographic information with space and aircraft remote sensing data and in managing large archives of remotely sensed data. SEIDAM will be composed of expert systems grouped in three levels. At the lowest level, the expert systems will manage and integrate data from diverse sources, taking account of symbolic representation differences and varying accuracies. Existing software can be controlled by these expert systems, without rewriting existing software into an Artificial Intelligence (AI) language. At the second level, SEIDAM will take the interpreted data (symbolic and numerical) and combine these with data models. At the top level, SEIDAM will respond to user goals for predictive outcomes given existing data. The SEIDAM Project will address the research areas of expert systems, data management, storage and retrieval, and user access and interfaces.

Goodenough, David G.↗

Why We Do What We Do: Data Reuse, Open Access and Privacy in Data Management at the Life Sciences Data Archive

As custodians of the unique and irreplaceable collections of human subject research data generated by the Human Research Program and its predecessors throughout the agency’s history, the Life Sciences Data Archive (LSDA) is charged with protecting participants’ privacy and implementing their consent decisions as it provides retrospective data for use in new studies. This active, stewardship-focused approach to data management and preservation shapes the products that LSDA provides to researchers and the responsibilities of researchers in using the data and publishing their results. This presentation reviews how federal and agency mandates shape LSDA’s data management procedures and expectations for researchers. Topics covered will include LSDA’s movement towards implementation of the FAIR (Findable, Accessible, Interoperable, Reusable) principles and the archive’s evolving data management practices; collaboration between LSDA and the Lifetime Surveillance of Astronaut Health (LSAH) project (the repository of astronaut medical data); LSDA’s response to the challenges of performing its stewardship role and maintaining trust given the public profiles of the subjects whose data it preserves; and the ever-increasing challenges to expectations of subject privacy stemming from the growing power and ubiquity of data analysis and aggregation tools.

data management↗

Why We Do What We Do: Data Reuse, Open Access, and Privacy in Data Management at the Life Sciences Data Archive

As custodian of the unique and irreplaceable collections of human subject research data generated by the Human Research Program and its predecessors throughout the agency’s history, the Life Sciences Data Archive (LSDA) is charged with protecting participants’ privacy and implementing their consent decisions as it provides retrospective data for use in new studies. This active, stewardship-focused approach to data management and preservation shapes the products that LSDA provides to researchers and the responsibilities of researchers in using the data and publishing their results. This presentation reviews how federal and agency mandates shape LSDA’s data management procedures and expectations for researchers. Topics covered will include LSDA’s movement towards implementation of the FAIR (Findable, Accessible, Interoperable, Reusable) principles and how the archive’s evolving data management practices support FAIR-ness; collaboration between LSDA and the Lifetime Surveillance of Astronaut Health (LSAH) project (the repository of astronaut medical data); LSDA’s response to the challenges of performing its stewardship role and maintaining trust given the public profiles of the subjects whose data it preserves; and the ever-increasing challenges to expectations of subject privacy stemming from the growing power and ubiquity of of data analysis and aggregation tools.

Data↗

A distributed data base management facility for the CAD/CAM environment

Current/PAD research in the area of distributed data base management considers facilities for supporting CAD/CAM data management in a heterogeneous network of computers encompassing multiple data base managers supporting a variety of data models. These facilities include coordinated execution of multiple DBMSs to provide for administration of and access to data distributed across them.

Balza, R. M.↗

The IsoGenie database: an interdisciplinary data management solution for ecosystems biology and environmental research

Modern microbial and ecosystem sciences require diverse interdisciplinary teams that are often challenged in “speaking” to one another due to different languages and data product types. Here we introduce the IsoGenie Database, a de novo developed data management and exploration platform, as a solution to this challenge of accurately representing and integrating heterogenous environmental and microbial data across ecosystem scales. The IsoGenieDB is a public and private data infrastructure designed to store and query data generated by the IsoGenie Project, a ~10 year DOE-funded project focused on discovering ecosystem climate feedbacks in a thawing permafrost landscape. The IsoGenieDB provides (i) a platform for IsoGenie Project members to explore the project’s interdisciplinary datasets across scales through the inherent relationships among data entities, (ii) a framework to consolidate and harmonize the datasets needed by the team’s modelers, and (iii) a public venue that leverages the same spatially explicit, disciplinarily integrated data structure to share published datasets. The IsoGenieDB is also being expanded to cover the NASA-funded Archaea to Atmosphere (A2A) project, which scales the findings of IsoGenie to a broader suite of Arctic peatlands, via the umbrella A2A Database (A2A-DB). The IsoGenieDB’s expandability and flexible architecture allow it to serve as an example ecosystems database.

54 ENVIRONMENTAL SCIENCES↗

Microbiome data management in action workshop: Atlanta, GA, USA, June 12–13, 2024

Microbiome research is revolutionizing human and environmental health, but the value and reuse of microbiome data are significantly hampered by the limited development and adoption of data standards. While several ongoing efforts are aimed at improving microbiome data management, significant gaps still remain in terms of defining and promoting adoption of consensus standards for these datasets. The Strengthening the Organization and Reporting of Microbiome Studies (STORMS) guidelines for human microbiome research have been endorsed and successfully utilized by many research organizations, publishers, and funding agencies, and have been recognized as a consensus community standard. No equivalent effort has occurred for environmental, synthetic, and non-human host-associated microbiomes. To address this growing need within the microbiome research community, we convened the Microbiome Data Management in Action Workshop (June 12–13, 2024, in Atlanta, GA, USA), to bring together key decision makers in microbiome science including researchers, publishers, funders, and data repositories. The 50 attendees, representing the diverse and interdisciplinary nature of microbiome research, discussed recent progress and challenges, and brainstormed actionable recommendations and paths forward for coordinated environmental microbiome data management and the modifications necessary for the STORMS guidelines to be applied to environmental, non-human host, and synthetic microbiomes. The outcomes of this workshop will form the basis of a formalized data management roadmap to be implemented across the field. These best practices will drive scientific innovation now and in years to come as these data continue to be used not only in targeted reanalyses but in large-scale models and machine learning efforts.

54 ENVIRONMENTAL SCIENCES↗

IPAD 2: Advances in Distributed Data Base Management for CAD/CAM

The Integrated Programs for Aerospace-Vehicle Design (IPAD) Project objective is to improve engineering productivity through better use of computer-aided design and manufacturing (CAD/CAM) technology. The focus is on development of technology and associated software for integrated company-wide management of engineering information. The objectives of this conference are as follows: to provide a greater awareness of the critical need by U.S. industry for advancements in distributed CAD/CAM data management capability; to present industry experiences and current and planned research in distributed data base management; and to summarize IPAD data management contributions and their impact on U.S. industry and computer hardware and software vendors.

Bostic, S. W.↗

Geospatial Data Platform for All

Spatiotemporal data has evolved in scale due to augmented use in cross-domain applications. Simultaneously, there is substantial growth in the availability of Geographic Information Systems (GIS) data provided by the United States Geological Survey (USGS) along with other federal, state, county, or local agencies through open-data portals and public access APIs. However, data availability does not equate with accessibility. Large-scale analyses and applications require robust, performant data management with co-location of data storage and computing. The insufficiency of data management infrastructure compels researchers to adopt ad hoc project- specific GIS data storage solutions (e.g., copying data to High-Performance computer file systems). As an ad hoc storage strategy does not scale, it hampers cross-domain analyses causing difficulty in data reuse and utilizing existing code bases. Furthermore, GIS data is complex and requires expertise to analyze and manipulate due to its intricate data structures and data-specific projection transformations. Despite the challenges, we recognize that derived GIS data products, e.g., satellite or LIDAR-based images, can be used in downstream applications such as AI by domain, but non-GIS experts. To address the data needs and overcome the challenges, we are working towards a GIS Data Platform focused on efficient data storage, data discovery and access, and an API to enable common workflows. We propose a knowledge-graph (KG) approach for data discovery, whereby datasets are semantically linked to higher- level constructs such as projects and research areas. The semantic data links enable researchers to explore datasets in a top-down approach by specifying relevant and meaningful terms (assists in finding hidden data). An advantage is that the nodes and edges in a knowledge graph create built-in semantic documentation. Deeper spatiotemporal connections between data sources can be encoded via Graph Neural Networks (GNN) (Zhang et al., 2021). The KG approach can be extended to integrate the data itself in a Virtual KG (VKG). Our work will derive inspiration from large-scale VKG efforts that have been undertaken or are currently underway as part of the OpenStreetMap project (Ding et al., 2021). For DOE Data Days, we share the proposed geospatial data platform hybrid (cloud/on-prem) architecture, our work-to-date on storing, retrieving, and transforming LiDAR and raster data relevant to two important NREL use-cases, including the Renewable Energy Potential (reV) Model, and present our proposal for a KG based data discovery engine.

data platform↗

A Lakehouse Architecture for the Management and Analysis of Heterogeneous Data for Biomedical Research and Mega-biobanks

Data Lakehouse is a new paradigm in data architectures that embodies and integrates already established concepts for the systematic management of disparate, large-scale data – a data lake for heterogeneous data management, use of open standards for high-performance querying, and systematic maintenance of the data "freshness". In addition to being a new concept, the data lakehouse is also still a conceptual construct. Many projects that use the lakehouse require maturing, empirical studies, and specific implementations. In this paper, we present our implementation of the data lakehouse concept in a biomedical research and health data analytics domain, and we discuss the implementation of some unique and novel features such as support for specialized access controls in support of HIPAA regulation and IRB protocols, and support for the FAIR standard.

Begoli, Edmon↗