Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Scientific data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Quality Control of and Analysis Enabling Use of MARCUS and MICRE data for Scientific Applications

The primary objective of this project was to provide a preliminary investigation of data collected during the MARCUS/MICRE field experiments and to produce products and analysis for the scientific community that will enable future scientific investigations and hypothesis testing. Further, the goal was to assist those in the scientific community to understand the strengths and caveats of the MARCUS/MICRE data so that members of the scientific community could use the data in their publications.

54 ENVIRONMENTAL SCIENCES↗

Automated system for measurement, collection and processing of hydrometeorological data aboard scientific research vessels of the GUGMS (SIGMA-s)

A report is made on the automated system known as SIGMA-s for the measurement, collection, and processing of hydrometeorological data aboard scientific research vessels of the Hydrometeorological Service. The various components of the system and the interfacing between them are described, as well as the projects that the system is equipped to handle.

Borisenkov, Y. P.↗

Apollo scientific experiments data handbook

A brief description of each of the Apollo scientific experiments was described, together with its operational history, the data content and formats, and the availability of the data. The lunar surface experiments described are the passive seismic, active seismic, lunar surface magnetometer, solar wind spectrometer, suprathermal ion detector, heat flow, charged particle, cold cathode gage, lunar geology, laser ranging retroreflector, cosmic ray detector, lunar portable magnetometer, traverse gravimeter, soil mechanics, far UV camera (lunar surface), lunar ejecta and meteorites, surface electrical properties, lunar atmospheric composition, lunar surface gravimeter, lunar seismic profiling, neutron flux, and dust detector. The orbital experiments described are the gamma-ray spectrometer, X-ray fluorescence, alpha-particle spectrometer, S-band transponder, mass spectrometer, far UV spectrometer, bistatic radar, IR scanning radiometer, particle shadows, magnetometer, lunar sounder, and laser altimeter. A brief listing of the mapping products available and information on the sample program were also included.

Eichelman, W. F.↗

NASA IKONOS Radiometric Characterization

NASA acquired imagery from the IKONOS satellite as part of its Scientific Data Purchase (SDP) program, which purchases scientific data sets from commercial sources. This viewgraph presentation describes the IKONOS satellite and its sensors, and then gives an overview of characterization efforts undertaken by NASA in cooperation with other government agencies. The characterization included relative radiometric correction, absolute radiometric characterization of data from Lunar Lake Playa, Nevada, and calibration of data from Stennis Space Center, Mississippi.

Pagnutti, Mary↗

Analysis of CrIS ATMS and AIRS AMSU Data Using Scientifically Equivalent Retrieval Algorithms

Monthly mean August 2014 Version-6.28 AIRS and CrIS products agree well with OMPS and CERES, and reasonably well with each other. Version-6.28 CrIS total precipitable water is biased dry compared to AIRS. AIRS and CrIS Version-6.36 water vapor products are both improved compared to Version-6.28. Version-6.36 AIRS and CrIS total precipitable water also shows improved agreement with each other. AIRS Version-6.36 total ozone agrees even better with OMPS than does AIRS Version-6.28, and gives reasonable results during polar winter where OMPS does not generate products. CrIS and ATMS are high spectral resolution IR and Microwave atmospheric sounders currently flying on the SNPP satellite, and are also scheduled for flight on future NPOESS satellites. CrIS/ATMS have similar sounding capabilities to those of the AIRS/AMSU sounder suite flying on EOS Aqua. The objective of this research is to develop and implement scientifically equivalent AIRS/AMSU and CrIS/ATMS retrieval algorithms with the goal of generating a continuous data record of AIRS/AMSU and CrIS/ATMS level-3 data products with a seamless transition between them in time. To achieve this, monthly mean AIRS/AMSU and CrIS/ATMS retrieved products, and more importantly their interannual differences, should show excellent agreement with each other. The currently operational AIRS Science Team Version-6 retrieval algorithm has generated 14 years of level-3 data products. A scientifically improved AIRS Version-7 retrieval algorithm is expected to become operational in 2017. We see significant improvements in water vapor and ozone in Version-7 retrieval methodology compared to Version-6.We are working toward finalization and implementation of scientifically equivalent AIRS/AMSU and CrIS/ATMS Version-7 retrieval algorithms to be used for the eventual processing of all AIRS/AMSU and CrIS/ATMS data. The latest version of our retrieval algorithm is Verison-6.36, which includes almost all the improvements we want in Version-7. Version-6.28 has been used to process both AIRS and CrIS data for August 2014. This poster compares August 2014 monthly mean Version-6.28 AIRS/AMSU and CrIS/ATMS products with each other, and also with monthly mean products obtained using AIRS Version-6. AIRS and CrIS results using Version-6.36 are presented for April 15, 2016. These demonstrate further improvements since Version-6.28. The new results also show improved agreement of Version-6.36 AIRS and CrIS products with each other. Version-6.36 is not yet optimized for CrIS ozone products.

AIRS↗

Publication of science data on CD-ROM: A guide and example

CD-ROM (Compact Disk-Read Only Memory) is becoming the standard media not only in audio recording, but also in the publication of data and information accessible on many computer platforms. Little has been written about the complicated process involved in creating easy-to-use, high quality, and useful CD-ROM's containing scientific data. This document is a manual designed to aid those who are responsible for the publication of scientific data on CD-ROM. All aspects and steps of the procedure are covered, from feasibility assessment through disk design, data preparation, disc mastering, and CD-ROM distribution. General advice and actual examples are based on lessons learned from the publication of scientific data for an interdisciplinary field experiment. Appendices include actual files from a CD-ROM, a purchase request for CD-ROM mastering services, and the disk art for the first disk published for the project.

Angelici, Gary↗

Data Readiness for Scientific AI at Scale

This paper examines how Data Readiness for AI (DRAI) principles apply to leadership-scale scientific datasets used to train foundation models. We analyze archetypal workflows across four representative domains—climate, nuclear fusion, bio/health, and materials—to identify common preprocessing patterns and domain-specific constraints. We introduce a two-dimensional readiness framework that combines canonical preprocessing patterns with a five-level operational readiness scale, both tailored to high-performance computing (HPC) environments. This framework helps outline key challenges in transforming large-scale scientific data into formats suitable for scalable AI training. Together, these dimensions form a conceptual maturity matrix that characterizes scientific data readiness and guides infrastructure development toward standardized, cross-domain support for scalable and reproducible AI for science.

Brewer, Wes [ORNL] (ORCID:0000000236393956)↗

Aligning NASA Earth Science Data Stewardship with FAIR Principles: Outcomes, Recommendations, and Future Directions

The FAIR Principles—Findable, Accessible, Interoperable, and Reusable—offer a widely accepted framework for improving the sharing and reuse of digital scientific data by both human and machine users. Following these principles is critical for effective scientific data stewardship, broader scientific collaboration, and compliance with federal and agency data policies. This paper, based on the work of NASA’s Open, Free, and FAIR Working Group (O’FAIR WG) under the Earth Science Data Systems Program, presents an overview of how FAIR is being applied within NASA’s Earth science data landscape. It highlights ongoing progress and challenges, identifies FAIR-enabling resources, and offers recommendations and strategic actions to enhance the FAIRness of NASA-funded open and free Earth science data products. The FAIR-enabling resources identified underscore the vital role of NASA's existing enterprise processes, standards, tools, and infrastructures in supporting FAIR implementation. Our findings show strong performance in making NASA Earth science data more findable and accessible. However, further work is needed—especially in enhancing interoperability, so that different systems and tools can better understand and exchange data. This is especially important for enabling machine-driven discovery and analysis. We emphasize the importance of a balanced strategy that combines a centralized, top-down approach—focused on building enterprise-level capabilities and processes—with a decentralized, bottom-up approach driven by discipline-specific needs and community practices. We advocate for coordinated efforts to enhance (meta)data interoperability to facilitate seamless data and information sharing and exchange of Earth science data both within NASA and across other agencies managing Earth science data.

Data Product↗

Extracting Material Property Measurement Data from Scientific Articles

Machine learning-based prediction of material properties is often hampered by the lack of sufficiently large training datasets. The majority of such measurement data is embedded in scientific literature and the ability to automatically extract these data is essential to support the development of reliable property prediction methods. In this work, we describe a methodology for an automatic property extraction framework using material solubility as the target property. We create an annotated dataset containing tags for solubility-related entities using a combination of regular expressions and manual tagging. We then compare five entity recognition models leveraging both token-level and span-level architectures on the task of classifying solute names, solubility values, and solubility units. Additionally, we explore a novel pretraining approach that leverages automated chemical name and quantity extraction tools to generate large datasets that do not rely on intensive manual effort. Finally, we perform an analysis to identify the causes of classification errors.

Panapitiya, Gihan U.↗

A History of NASA Remote Sensing Contributions to Archaeology

During its long history of developing and deploying remote sensing instruments, NASA has provided a scientific data that have benefitted a variety of scientific applications among them archaeology. Multispectral and hyperspectral instrument mounted on orbiting and suborbital platforms have provided new and important information for the discovery, delineation and analysis of archaeological sites worldwide. Since the early 1970s, several of the ten NASA centers have collaborated with archaeologists to refine and validate the use of active and passive remote sensing for archeological use. The Stennis Space Center (SSC), located in Mississippi USA has been the NASA leader in archeological research. Together with colleagues from Goddard Space Flight Center (GSFC), Marshall Space Flight Center (MSFC), and the Jet Propulsion Laboratory (JPL), SSC scientists have provided the archaeological community with useful images and sophisticated processing that have pushed the technological frontiers of archaeological research and applications. Successful projects include identifying prehistoric roads in Chaco canyon, identifying sites from the Lewis and Clark Corps of Discovery exploration and assessing prehistoric settlement patterns in southeast Louisiana. The Scientific Data Purchase (SDP) stimulated commercial companies to collect archaeological data. At present, NASA formally solicits "space archaeology" proposals through its Earth Science Directorate and continues to assist archaeologists and cultural resource managers in doing their work more efficiently and effectively. This paper focuses on passive remote sensing and does not consider the significant contributions made by NASA active sensors. Hyperspectral data offers new opportunities for future archeological discoveries.

Giardino, Marco J.↗

Visualization Quality Assessment

Understanding how inaccuracies in visualizations affect users’ perception and understanding of scientific data is hard. Inaccuracies in visualizations are quite common and could arise from a range of sources such as errors in the original dataset arising from compression artifacts, errors in the capturing device, noise during transmission of the data, effects due to the algorithm being used to convert data to visualization images, images generated from neural networks, and sources we have yet to discover. Many image quality assessment metrics have been developed to quantify image errors. However, these are usually focused on “natural images” rather than visualizations of scientific data. Common image quality assessment metrics (IQAs) include MSE, PSNR, perceptual metrics such SSIM, FSIM as well as perceptual metrics using deep learning approaches. However, a critical part of understanding how errors are perceived by humans, and subsequently developing more accurate quality assessment metrics, is through user evaluation studies. The goal of this software is to develop a visualization quality assessment (VQA) process that will enable the generation of VQAs that can be used to quantify errors in scientific data visualizations. The VQA development process will include software to support user evaluation experimental design, analysis of visualization differences against standard quality metrics, and the ability to develop additional VQA metrics specific to scientific visualization images.

Grosset, Andre↗

Tool for Automated Retrieval of Generic Event Tracks (TARGET)

Methods have been developed to identify and track tornado-producing mesoscale convective systems (MCSs) automatically over the continental United States, in order to facilitate systematic studies of these powerful and often destructive events. Several data sources were combined to ensure event identification accuracy. Records of watches and warnings issued by National Weather Service (NWS), and tornado locations and tracks from the Tornado History Project (THP) were used to locate MCSs in high-resolution precipitation observations and GOES infrared (11-micron) Rapid Scan Operation (RSO) imagery. Thresholds are then applied to the latter two data sets to define MCS events and track their developments. MCSs produce a broad range of severe convective weather events that are significantly affecting the living conditions of the populations exposed to them. Understanding how MCSs grow and develop could help scientists improve their weather prediction models, and also provide tools to decision-makers whose goals are to protect populations and their property. Associating storm cells across frames of remotely sensed images poses a difficult problem because storms evolve, split, and merge. Any storm-tracking method should include the following processes: storm identification, storm tracking, and quantification of storm intensity and activity. The spatiotemporal coordinates of the tracks will enable researchers to obtain other coincident observations to conduct more thorough studies of these events. In addition to their tracked locations, their areal extents, precipitation intensities, and accumulations all as functions of their evolutions in time were also obtained and recorded for these events. All parameters so derived can be catalogued into a moving object database (MODB) for custom queries. The purpose of this software is to provide a generalized, cross-platform, pluggable tool for identifying events within a set of scientific data based upon specified criteria with the possibility of storing identified events into a searchable database. The core of the application uses an implementation of the connected component labeling (CCL) algorithm to identify areas of interest, then uses a set of criteria to establish spatial and temporal relationships between identified components. The CCL algorithm is used for identifying objects within images for computer vision. This application applies it to scientific data sets using arbitrary criteria. The most novel concept was applying a generalized CCL implementation to scientific data sets for establishing events both spatially and temporally. The combination of several existing concepts (pluggable components, generalized CCL algorithm, etc.) into one application is also novel. In addition, how the system is designed, i.e., its extensibility with pluggable components, and its configurability with a simple configuration file, is innovative. This allows the system to be applied to new scenarios with ease.

Clune, Thomas↗

The Role of Advanced Information System Technology in Remote Sensing for NASA's Earth Science Enterprise in the 21st Century

Future NASA Earth observing satellites will carry high-precision instruments capable of producing large amounts of scientific data. The strategy will be to network these instrument-laden satellites into a web-like array of sensors to facilitate the collection, processing, transmission, storage, and distribution of data and data products - the essential elements of what we refer to as "Information Technology." Many of these Information Technologies will enable the satellite and ground information systems to function effectively in real-time, providing scientists with the capability of customizing data collection activities on a satellite or group of satellites directly from the ground. In future systems, extremely large quantities of data collected by scientific instruments will require the fastest processors, the highest communication channel transfer rates, and the largest data storage capacity to insure that data flows smoothly from the satellite-based instrument to the ground-based archive. Autonomous systems will control all essential processes and play a key role in coordinating the data flow through space-based communication networks. In this paper, we will discuss those critical information technologies for Earth observing satellites that will support the next generation of space-based scientific measurements of planet Earth, and insure that data and data products provided by these systems will be accessible to scientists and the user community in general.

Prescott, Glenn↗

Generic functional requirements for a NASA general-purpose data base management system

Generic functional requirements for a general-purpose, multi-mission data base management system (DBMS) for application to remotely sensed scientific data bases are detailed. The motivation for utilizing DBMS technology in this environment is explained. The major requirements include: (1) a DBMS for scientific observational data; (2) a multi-mission capability; (3) user-friendly; (4) extensive and integrated information about data; (5) robust languages for defining data structures and formats; (6) scientific data types and structures; (7) flexible physical access mechanisms; (8) ways of representing spatial relationships; (9) a high level nonprocedural interactive query and data manipulation language; (10) data base maintenance utilities; (11) high rate input/output and large data volume storage; and adaptability to a distributed data base and/or data base machine configuration. Detailed functions are specified in a top-down hierarchic fashion. Implementation, performance, and support requirements are also given.

Lohman, G. M.↗

Improve Data Mining and Knowledge Discovery Through the Use of MatLab

Data mining is widely used to mine business, engineering, and scientific data. Data mining uses pattern based queries, searches, or other analyses of one or more electronic databases/datasets in order to discover or locate a predictive pattern or anomaly indicative of system failure, criminal or terrorist activity, etc. There are various algorithms, techniques and methods used to mine data; including neural networks, genetic algorithms, decision trees, nearest neighbor method, rule induction association analysis, slice and dice, segmentation, and clustering. These algorithms, techniques and methods used to detect patterns in a dataset, have been used in the development of numerous open source and commercially available products and technology for data mining. Data mining is best realized when latent information in a large quantity of data stored is discovered. No one technique solves all data mining problems; challenges are to select algorithms or methods appropriate to strengthen data/text mining and trending within given datasets. In recent years, throughout industry, academia and government agencies, thousands of data systems have been designed and tailored to serve specific engineering and business needs. Many of these systems use databases with relational algebra and structured query language to categorize and retrieve data. In these systems, data analyses are limited and require prior explicit knowledge of metadata and database relations; lacking exploratory data mining and discoveries of latent information. This presentation introduces MatLab(R) (MATrix LABoratory), an engineering and scientific data analyses tool to perform data mining. MatLab was originally intended to perform purely numerical calculations (a glorified calculator). Now, in addition to having hundreds of mathematical functions, it is a programming language with hundreds built in standard functions and numerous available toolboxes. MatLab's ease of data processing, visualization and its enormous availability of built in functionalities and toolboxes make it suitable to perform numerical computations and simulations as well as a data mining tool. Engineers and scientists can take advantage of the readily available functions/toolboxes to gain wider insight in their perspective data mining experiments.

Shaykhian, Gholam Ali↗

Improve Data Mining and Knowledge Discovery through the use of MatLab

Data mining is widely used to mine business, engineering, and scientific data. Data mining uses pattern based queries, searches, or other analyses of one or more electronic databases/datasets in order to discover or locate a predictive pattern or anomaly indicative of system failure, criminal or terrorist activity, etc. There are various algorithms, techniques and methods used to mine data; including neural networks, genetic algorithms, decision trees, nearest neighbor method, rule induction association analysis, slice and dice, segmentation, and clustering. These algorithms, techniques and methods used to detect patterns in a dataset, have been used in the development of numerous open source and commercially available products and technology for data mining. Data mining is best realized when latent information in a large quantity of data stored is discovered. No one technique solves all data mining problems; challenges are to select algorithms or methods appropriate to strengthen data/text mining and trending within given datasets. In recent years, throughout industry, academia and government agencies, thousands of data systems have been designed and tailored to serve specific engineering and business needs. Many of these systems use databases with relational algebra and structured query language to categorize and retrieve data. In these systems, data analyses are limited and require prior explicit knowledge of metadata and database relations; lacking exploratory data mining and discoveries of latent information. This presentation introduces MatLab(TradeMark)(MATrix LABoratory), an engineering and scientific data analyses tool to perform data mining. MatLab was originally intended to perform purely numerical calculations (a glorified calculator). Now, in addition to having hundreds of mathematical functions, it is a programming language with hundreds built in standard functions and numerous available toolboxes. MatLab's ease of data processing, visualization and its enormous availability of built in functionalities and toolboxes make it suitable to perform numerical computations and simulations as well as a data mining tool. Engineers and scientists can take advantage of the readily available functions/toolboxes to gain wider insight in their perspective data mining experiments.

Shaykahian, Gholan Ali↗