Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “scientific data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Performance and Reliability Assessment of the U.S. Department of Energy Atmospheric Radiation Measurement (ARM) Data Advisor (ADA)

The Atmospheric Radiation Measurement (ARM) User Facility provides one of the world's largest openly accessible repositories of atmospheric observations through the ARM Data Discovery platform. Although the repository contains more than three decades of measurements collected from permanent observatories, mobile facilities, aircraft campaigns, and field experiments, identifying appropriate datasets can be challenging, particularly for new users unfamiliar with ARM instrumentation and datastream organization. To improve data accessibility, the ARM Data Center developed the ARM Data Advisor (ADA), an artificial intelligence-powered assistant designed to facilitate scientific data discovery, dataset interpretation, and user guidance. This report evaluates ADA's performance as a domain-specific scientific assistant using realistic atmospheric science workflows. The evaluation examines five key capabilities: data retrieval and curation efficiency, hallucination resistance, scientific reasoning, response to ambiguous queries, and content retention and session continuity. Representative prompts were developed to simulate typical interactions between researchers and the ARM Data Discovery platform, and ADA's responses were assessed for retrieval completeness, scientific accuracy, consistency, and practical usefulness. In these representative tests, ADA reduced the complexity of discovering and accessing ARM datasets by recommending appropriate datastreams, explaining instrumentation, interpreting metadata, and assisting with data processing workflows. ADA also exhibits strong domain knowledge of atmospheric science terminology and generally resists hallucination by acknowledging unavailable datasets and requesting clarification when appropriate. Overall, the results indicate that ADA represents a promising advancement in scientific data discovery within the ARM User Facility and has considerable potential to improve researcher productivity, particularly for new users and interdisciplinary scientists seeking efficient access to ARM observations.

Salvador, Christian [ORNL] (ORCID:0000000283287777↗

Usable Data Abstractions for Next-Generation Scientific Workflows

Data- and computationally-intensive scientific research, such as numerical simulations and inversions or the training of large neural networks in machine learning applications, that are well suited for HPC environments also often require expert insight and evaluation throughout the computation which can be greatly facilitated with the use of interactive computing tools, such as those in the Jupyter ecosystem. HPC workflows and interactive workflows are typically treated as orthogonal, however, the next generation of research will require both. The first challenge we face in this project is thus designing the right level of abstractions to allow interactive capabilities in the JupyterLab environment to allow the working scientist to flexibly explore and query their data at multiple levels, with a minimal amount of customization required of the underlying optimized codes. In addition to these questions regarding the high-level representation of data for interactive use in HPC, we tackled two additional issues that are part of the entire lifecycle of research and that become particularly acute in HPC contexts: how to improve the experience of interfacing with the HPC system's scheduling environment for a scientist focused on exploratory questions, and how can that scientist then best share the results of their work with others in a self-contained, reproducible manner.

97 MATHEMATICS AND COMPUTING↗

Science Goal Driven Automation for NASA Missions: The Science Goal Monitor

Infusion of automation technologies into NASA s future missions will be essential not only to achieve substantial reduction in mission operations staff and costs, but also in order to both effectively handle an exponentially increasing volume of scientific data and to successfully meet dynamic, opportunistic scientific goals and objectives. Current spacecraft operations cannot respond to science driven events, such as intrinsically variable or short-lived phenomena in a timely manner. For such investigations, we must teach our platforms to dynamically understand, recognize, and react to the scientists goals. While much effort has gone into automating routine spacecraft operations to reduce human workload and hence costs, applying intelligent automation to the science side, i.e., science data acquisition, data analysis and reactions to that data analysis in a timely and still scientifically valid manner, has been relatively under-emphasized.

Korathkar, Anuradha↗

A data management system for engineering and scientific computing

Data elements and relationship definition capabilities for this data management system are explicitly tailored to the needs of engineering and scientific computing. System design was based upon studies of data management problems currently being handled through explicit programming. The system-defined data element types include real scalar numbers, vectors, arrays and special classes of arrays such as sparse arrays and triangular arrays. The data model is hierarchical (tree structured). Multiple views of data are provided at two levels. Subschemas provide multiple structural views of the total data base and multiple mappings for individual record types are supported through the use of a REDEFINES capability. The data definition language and the data manipulation language are designed as extensions to FORTRAN. Examples of the coding of real problems taken from existing practice in the data definition language and the data manipulation language are given.

Elliot, L.↗

Catalog of Viking mission data

This catalog announces the present/expected availability of scientific data acquired by the Viking missions and contains descriptions of the Viking spacecraft, experiments, and data sets. An index is included listing the team leaders and team members for the experiments. Information on NSSDC facilities and ordering procedures, and a list of acronyms and abbreviations are included in the appendices.

Vostreys, R. W.↗

Uniform-in-phase-space data selection with iterative normalizing flows

Improvements in computational and experimental capabilities are rapidly increasing the amount of scientific data that are routinely generated. In applications that are constrained by memory and computational intensity, excessively large datasets may hinder scientific discovery, making data reduction a critical component of data-driven methods. Datasets are growing in two directions: the number of data points and their dimensionality. Whereas dimension reduction typically aims at describing each data sample on lower-dimensional space, the focus here is on reducing the number of data points. A strategy is proposed to select data points such that they uniformly span the phase-space of the data. The algorithm proposed relies on estimating the probability map of the data and using it to construct an acceptance probability. An iterative method is used to accurately estimate the probability of the rare data points when only a small subset of the dataset is used to construct the probability map. Instead of binning the phase-space to estimate the probability map, its functional form is approximated with a normalizing flow. Therefore, the method naturally extends to high-dimensional datasets. The proposed framework is demonstrated as a viable pathway to enable data-efficient machine learning when abundant data are available.

97 MATHEMATICS AND COMPUTING↗

Tracking Provenance of Earth Science Data

Tremendous volumes of data have been captured, archived and analyzed. Sensors, algorithms and processing systems for transforming and analyzing the data are evolving over time. Web Portals and Services can create transient data sets on-demand. Data are transferred from organization to organization with additional transformations at every stage. Provenance in this context refers to the source of data and a record of the process that led to its current state. It encompasses the documentation of a variety of artifacts related to particular data. Provenance is important for understanding and using scientific datasets, and critical for independent confirmation of scientific results. Managing provenance throughout scientific data processing has gained interest lately and there are a variety of approaches. Large scale scientific datasets consisting of thousands to millions of individual data files and processes offer particular challenges. This paper uses the analogy of art history provenance to explore some of the concerns of applying provenance tracking to earth science data. It also illustrates some of the provenance issues with examples drawn from the Ozone Monitoring Instrument (OMI) Data Processing System (OMIDAPS) run at NASA's Goddard Space Flight Center by the first author.

Tilmes, Curt↗

Hierarchical Analysis of Halo Center in Cosmology

Ever-increasing data size raises many challenges for scientific data analysis. Particularly in cosmological N-body simulation, finding the center of a dark matter halo suffers heavily from the large computational cost associated with the large number of particles (up to 20 million). In this work, we exploit the latent structure embed in a halo, and we propose a hierarchical approach to approximate the exact gravitational potential calculation for each particle in order to more efficiently find the halo center. Tests of our method on data from N-body simulations show that in many cases the hierarchical algorithm performs significantly faster than existing methods with a desirable accuracy.

97 MATHEMATICS AND COMPUTING↗

Balloon stratospheric research flights, November 1974 to January 1976

These flights were designed to measure the vertical concentration profile of trace stratospheric species which form major links in the photochemical system of the upper atmosphere. An overview of the specific goals of the program, a statement of program management and support functions, a brief description of the instrumentation flown, pertinent engineering and payload operations data, and a summary of the scientific data obtained for each of the last five flights during this period are presented.

Allen, N. C.↗

Payload data processing for the Space Shuttle program

The paper examines communications and data handling services that the Space Shuttle will provide for payloads and discusses mission and data processing capabilities. Uses of these capabilities are considered, and the Orbiter data processing system is described. Functional scientific data interfaces for attached payloads and functional system status data interfaces for attached payloads are indicated.

Batson, B. H.↗

Data access for scientific problem solving

An essential ingredient in scientific work is data. In disciplines such as Oceanography, data sources are many and volumes are formidable. The full value of large stores of data cannot be realized unless careful thought is given to data access. JPL has developed the Pilot Ocean Data System to investigate techniques for archiving and accessing ocean data obtained from space. These include efficient storage and rapid retrieval of satellite data, an easy-to-use user interface, and a variety of output products which, taken together, permit researchers to extract and use data rapidly and conveniently.

Brown, James W.↗

The df: A proposed data format standard

A standard is proposed describing a portable format for electronic exchange of data in the physical sciences. Writing scientific data in a standard format has three basic advantages: portability; the ability to use metadata to aid in interpretation of the data (understandability); and reusability. An improperly formulated standard format tends towards four disadvantages: (1) it can be inflexible and fail to allow the user to express his data as needed; (2) reading and writing such datasets can involve high overhead in computing time and storage space; (3) the format may be accessible only on certain machines using certain languages; and (4) under some circumstances it may be uncertain whether a given dataset actually conforms to the standard. A format was designed which enhances these advantages and lessens the disadvantages. The fundamental approach is to allow the user to make her own choices regarding strategic tradeoffs to achieve the performance desired in her local environment. The choices made are encoded in a specific and portable way in a set of records. A fully detailed description and specification of the format is given, and examples are used to illustrate various concepts. Implementation is discussed.

Lait, Leslie R.↗

Visualizing Parallel Computer System Performance

Parallel computer systems are among the most complex of man's creations, making satisfactory performance characterization difficult. Despite this complexity, there are strong, indeed, almost irresistible, incentives to quantify parallel system performance using a single metric. The fallacy lies in succumbing to such temptations. A complete performance characterization requires not only an analysis of the system's constituent levels, it also requires both static and dynamic characterizations. Static or average behavior analysis may mask transients that dramatically alter system performance. Although the human visual system is remarkedly adept at interpreting and identifying anomalies in false color data, the importance of dynamic, visual scientific data presentation has only recently been recognized Large, complex parallel system pose equally vexing performance interpretation problems. Data from hardware and software performance monitors must be presented in ways that emphasize important events while eluding irrelevant details. Design approaches and tools for performance visualization are the subject of this paper.

Malony, Allen D.↗

Project Genesis: Mars in situ propellant technology demonstrator mission

Project Genesis is a low cost, near-term, unmanned Mars mission, whose primary purpose is to demonstrate in situ resource utilization (ISRU) technology. The essence of the mission is to use indigenously produced fuel and oxidizer to propel a ballistic hopper. The Mars Landing Vehicle/Hopper (MLVH) has an Earth launch mass of 625 kg and is launched aboard a Delta 117925 launch vehicle into a conjunction class transfer orbit to Mars. Upon reaching its target, the vehicle performs an aerocapture maneuver and enters an elliptical orbit about Mars. Equipped with a ground penetrating radar, the MLVH searches for subsurface water ice deposits while in orbit for several weeks. A deorbit burn is then performed to bring the MLVH into the Martian atmosphere for landing. Following aerobraking and parachute deployment, the vehicle retrofires to a soft landing on Mars. Once on the surface, the MLVH begins to acquire scientific data and to manufacture methane and oxygen via the Sabatier process. This results in a fuel-rich O2/CH4 mass ratio of 2, which yields a sufficiently high specific impulse (335 sec) that no additional oxygen need be manufactured, thus greatly simplifying the design of the propellant production plant. During a period of 153 days the MLVH produces and stores enough fuel and oxidizer to make a 30 km ballistic hop to a different site of scientific interest. At this new location the MLVH resumes collecting surface and atmospheric data with the onboard instrumentation. Thus, the MLVH is able to provide a wealth of scientific data which would otherwise require two separate missions or separate vehicles, while proving a new and valuable technology that will facilitate future unmanned and manned exploration of Mars. Total mission cost, including the Delta launch vehicle, is estimated to be $200 million.

Acosta, Francisco Garcia↗

U.S. data processing for the IRAS project

The JPL's Scientific Data Analysis System (SDAS), which will process IRAS data and produce a catalogue of perhaps a million infrared sources in the sky, as well as other information for astronomical records, is described. The purposes of SDAS are discussed, and the major SDAS processors are shown in block diagram. The catalogue processing is addressed, mentioning the basic processing steps which will be applied to raw detector data. Signal reconstruction and conversion to astrophysical units, source detection, source confirmation, data management, and survey data products are considered in detail.

Duxbury, J. H.↗

An interactive environment for the analysis of large Earth observation and model data sets

We propose to develop an interactive environment for the analysis of large Earth science observation and model data sets. We will use a standard scientific data storage format and a large capacity (greater than 20 GB) optical disk system for data management; develop libraries for coordinate transformation and regridding of data sets; modify the NCSA X Image and X Data Slice software for typical Earth observation data sets by including map transformations and missing data handling; develop analysis tools for common mathematical and statistical operations; integrate the components described above into a system for the analysis and comparison of observations and model results; and distribute software and documentation to the scientific community.

Bowman, Kenneth P.↗

An interactive environment for the analysis of large Earth observation and model data sets

We propose to develop an interactive environment for the analysis of large Earth science observation and model data sets. We will use a standard scientific data storage format and a large capacity (greater than 20 GB) optical disk system for data management; develop libraries for coordinate transformation and regridding of data sets; modify the NCSA X Image and X DataSlice software for typical Earth observation data sets by including map transformations and missing data handling; develop analysis tools for common mathematical and statistical operations; integrate the components described above into a system for the analysis and comparison of observations and model results; and distribute software and documentation to the scientific community.

Bowman, Kenneth P.↗