Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “open data format”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

STELAR: An experiment in the electronic distribution of astronomical literature

STELAR (Study of Electronic Literature for Astronomical Research) is a Goddard-based project designed to test methods of delivering technical literature in machine readable form. To that end, we have scanned a five year span of the ApJ, ApJ Supp, AJ and PASP, and have obtained abstracts for eight leading academic journals from NASA/STI CASI, which also makes these abstracts available through the NASA RECON system. We have also obtained machine readable versions of some journal volumes from the publishers, although in many instances, the final typeset versions are no longer available. The fundamental data object for the STELAR database is the article, a collection of items associated with a scientific paper - abstract, scanned pages (in a variety of formats), figures, OCR extractions, forward and backward references, errata and versions of the paper in various formats (e.g., TEX, SGML, PostScript, DVI). Articles are uniquely referenced in the database by journal name, volume number and page number. The selection and delivery of articles is accomplished through the WAIS (Wide Area Information Server) client/server models requiring only an Internet connection. Modest modifications to the server code have made it capable of delivering the multiple data types required by STELAR. WAIS is a platform independent and fully open multi-disciplinary delivery system, originally developed by Thinking Machines Corp. and made available free of charge. It is based on the ISO Z39.50 standard communications protocol. WAIS servers run under both UNIX and VMS. WAIS clients run on a wide variety of machines, from UNIX-based Xwindows systems to MS-DOS and macintosh microcomputers. The WAIS system includes full-test indexing and searching of documents, network interface and easy access to a variety of document viewers. ASCII versions of the CASI abstracts have been formatted for display and the full test of the abstracts has been indexed. The entire WAIS database of abstracts is now available for use by the astronomical community. Enhancements of the search and retrieval system are under investigation to include specialized searches (by reference, author or keyword, as opposed to full test searches), improved handling of word stems, improvements in relevancy criteria and other retrieval techniques, such as factor spaces. The STELAR project has been assisted by the full cooperation of the AAS, the ASP, the publishers of the academic journals, librarians from GSFC, NRAO and STScI, the Library of Congress, and the University of North Carolina at Chapel Hill.

Warnock, A.↗

Potential Multi-Component Structure of the Debris Disk Around HIP 17439 Revealed by Herschel DUNES

Context. The dust observed in debris disks is produced through collisions of larger bodies left over from the planet/planetesimal formation process. Spatially resolving these disks permits to constrain their architecture and thus that of the underlying planetary/planetesimal system. Aims. Our Herschel open time key program DUNES aims at detecting and characterizing debris disks around nearby, sun-like stars. In addition to the statistical analysis of the data, the detailed study of single objects through spatially resolving the disk and detailed modeling of the data is a main goal of the project. Methods. We obtained the first observations spatially resolving the debris disk around the sun-like star HIP 17439 (HD 23484) using the instruments PACS and SPIRE on board the Herschel Space Observatory. Simultaneous multi-wavelength modeling of these data together with ancillary data from the literature is presented. Results. A standard single component disk model fails to reproduce the major axis radial profiles at 70 μm, 100 μm, and 160 μm simultaneously. Moreover, the best-fit parameters derived from such a model suggest a very broad disk extending from few au up to few hundreds of au from the star with a nearly constant surface density which seems physically unlikely. However, the constraints from both the data and our limited theoretical investigation are not strong enough to completely rule out this model. An alternative, more plausible, and better fitting model of the system consists of two rings of dust at approx. 30 au and 90 au, respectively, while the constraints on the parameters of this model are weak due to its complexity and intrinsic degeneracies. Conclusions. The disk is probably composed of at least two components with different spatial locations (but not necessarily detached), while a single, broad disk is possible, but less likely. The two spatially well-separated rings of dust in our best-fit model suggest the presence of at least one high mass planet or several low-mass planets clearing the region between the two rings from planetesimals and dust.

sun-like stars↗

User manual for NASA Lewis 10 by 10 foot supersonic wind tunnel

This manual describes the 10- by 10-Foot Supersonic Wind Tunnel at the NASA Lewis Research Center and provides information for users who wish to conduct experiments in this facility. Tunnel performance operating envelopes of altitude, dynamic pressure, Reynolds number, total pressure, and total temperature as a function of test section Mach number are presented. Operating envelopes are shown for both the aerodynamic (closed) cycle and the propulsion (open) cycle. The tunnel test section Mach number range is 2.0 to 3.5. General support systems, such as air systems, hydraulic system, hydrogen system, fuel system, and Schlieren system, are described. Instrumentation and data processing and acquisition systems are also described. Pretest meeting formats and schedules are outlined. Tunnel user responsibility and personnel safety are also discussed.

Soeder, Ronald H.↗

A Community Convention for Ecological Forecasting: Output Files and Metadata Version 1.0

This paper summarizes the open community conventions developed by the Ecological Forecasting Initiative (EFI) for the common formatting and archiving of ecological forecasts and the metadata associated with these forecasts. Such open standards are intended to promote interoperability and facilitate forecast communication, distribution, validation, and synthesis. For output files, we first describe the convention conceptually in terms of global attributes, forecast dimensions, forecasted variables, and ancillary indicator variables. We then illustrate the application of this convention to the two file formats that are currently preferred by the EFI, netCDF (network common data form), and comma-separated values (CSV), but note that the convention is extensible to future formats. For metadata, EFI's convention identifies a subset of conventional metadata variables that are required (e.g., temporal resolution and output variables) but focuses on developing a framework for storing information about forecast uncertainty propagation, data assimilation, and model complexity, which aims to facilitate cross-forecast synthesis. The initial application of this convention expands upon the Ecological Metadata Language (EML), a commonly used metadata standard in ecology. To facilitate community adoption, we also provide a Github repository containing a metadata validator tool and several vignettes in R and Python on how to both write and read in the EFI standard. Lastly, we provide guidance on forecast archiving, making an important distinction between short-term dissemination and long-term forecast archiving, while also touching on the archiving of code and workflows. Overall, the EFI convention is a living document that can continue to evolve over time through an open community process.

Michael C. Dietze↗

Steps Toward Improved Integration, Search, and Analysis of Heterogeneous Data in the Astrobiology Habitable Environments Database

The Astrobiology Habitable Environments Database (AHED) is a new data system being developed as a long-term, open-access repository for astrobiology data. AHED is intended to store user-contributed results from NASA or externally-funded research in astrobiology, and to encourage sharing and synergy within the astrobiology community. However, the interdisciplinary nature of astrobiology presents some specific challenges to data management, integration, and analysis within AHED. In some disciplines (e.g., genomics), open databases thrive because the contributed products are fairly uniform and standardized (e.g., sequence data). In astrobiology, each investigation produces a unique set of data products; this makes it difficult to search across different datasets to find similar data, or to combine results from separate investigations. With AHED, we are taking steps to ensure there is adequate metadata - both at the dataset and record levels - to facilitate search, integration, and analysis. At the dataset level, we are developing a new metadata standard for describing astrobiology datasets, with detailed information about content, funding source, and scientific relevance, along with a set of topical keywords for characterizing datasets. At the record level, we are encouraging users to provide more structured content and finer-grained metadata. In many user-contributed science data repositories, few restrictions are placed on the uploaded data format, and minimal or no record-level metadata is required; thus users are unburdened when it comes to data preparation. The tradeoff is that deep integration and search across datasets is almost impossible without standardized structures and metadata. Although AHED users are free to upload minimally-described datasets, they will be encouraged to use database authoring tools (supplied by the underlying platform - Open Data Repository's Data Publisher) plus a set of customizable astrobiology-specific templates to help structure their data and provide standardized metadata. In reward for their extra effort, AHED will be able to deliver enhanced search, discovery, and analysis capabilities.

astrobiology↗

CFD Extraction Tool for TecPlot From DPLR Solutions

This invention is a TecPlot macro of a computer program in the TecPlot programming language that processes data from DPLR solutions in TecPlot format. DPLR (Data-Parallel Line Relaxation) is a NASA computational fluid dynamics (CFD) code, and TecPlot is a commercial CFD post-processing tool. The Tec- Plot data is in SI units (same as DPLR output). The invention converts the SI units into British units. The macro modifies the TecPlot data with unit conversions, and adds some extra calculations. After unit conversions, the macro cuts a slice, and adds vectors on the current plot for output format. The macro can also process surface solutions. Existing solutions use manual conversion and superposition. The conversion is complicated because it must be applied to a range of inter-related scalars and vectors to describe a 2D or 3D flow field. It processes the CFD solution to create superposition/comparison of scalars and vectors. The existing manual solution is cumbersome, open to errors, slow, and cannot be inserted into an automated process. This invention is quick and easy to use, and can be inserted into an automated data-processing algorithm.

Norman, David↗

The European Southern Observatory-MIDAS table file system

The new and substantially upgraded version of the Table File System in MIDAS is presented as a scientific database system. MIDAS applications for performing database operations on tables are discussed, for instance, the exchange of the data to and from the TFS, the selection of objects, the uncertainty joins across tables, and the graphical representation of data. This upgraded version of the TFS is a full implementation of the binary table extension of the FITS format; in addition, it also supports arrays of strings. Different storage strategies for optimal access of very large data sets are implemented and are addressed in detail. As a simple relational database, the TFS may be used for the management of personal data files. This opens the way to intelligent pipeline processing of large amounts of data. One of the key features of the Table File System is to provide also an extensive set of tools for the analysis of the final results of a reduction process. Column operations using standard and special mathematical functions as well as statistical distributions can be carried out; commands for linear regression and model fitting using nonlinear least square methods and user-defined functions are available. Finally, statistical tests of hypothesis and multivariate methods can also operate on tables.

Peron, M.↗

Astronomical Data Center Bulletin, volume 1, number 2

Work in progress on astronomical catalogs is presented in 16 papers. Topics cover astronomical data center operations; automatic astronomical data retrieval at GSFC; interactive computer reference search of astronomical literature 1950-1976; formatting, checking, and documenting machine-readable catalogs; interactive catalog of UV, optical, and HI data for 201 Virgo cluster galaxies; machine-readable version of the general catalog of variable stars, third edition; galactic latitude and magnitude distribution of two astronomical catalogs; the catalog of open star clusters; infrared astronomical data base and catalog of infrared observations; the Air Force geophysics laboratory; revised magnetic tape of the N30 catalog of 5,268 standard stars; positional correlation of the two-micron sky survey and Smithsonian Astrophysical Observatory catalog sources; search capabilities for the catalog of stellar identifications (CSI) 1979 version; CSI statistics: blue magnitude versus spectral type; catalogs available from the Astronomical Data Center; and status report on machine-readable astronomical catalogs.

Nagy, T. A.↗

Optimum Detection of Frequency-Hopped Signals

This paper derives and analyzes optimum and near-optimum structures for detecting frequency-hopped (FH) signals with arbitrary modulation in additive white Gaussian noise. The principalmodulation formats considered are M-ary frequency-shift-keying (MFSK) with fast frequency hopping(FFH) wherein a single tone is transmitted per hop, and slow frequency hopping (SFH) with multipleMFSK tones (data symbols) per hop. The SFH detection category has not previously been addressedin the open literature and its analysis is generally more complex than FFH.

Cheng, Unjeng↗

A Concept for Civil Space Traffic Management

As technology has improved, operators have sought to use cubesats, as well as smallsats more generally, to perform increasingly more ambitious and sophisticated functions. Despite this, practical concerns associated with cubesat infant mortality, conjunctions, limited maneuverability, and debris generation have been relatively muted because most cubesats have been launched to lower orbits that limit both their orbital lifetime and consequences should a collision occur. NASA ARC has developed a concept for a highly-automated and distributed space traffic management (STM) architecture, drawing on similar work done to provide traffic management for small unmanned aerial systems (UAS) operating at low altitudes. The system proposes a strategy to accommodate growing space traffic volume safely, as well as pave the way for a transition of civil STM authority to a civilian governmental entity. The architecture envisions an open-access software platform architecture of data and service suppliers, consumers, and regulators, connected via a set of application programming interfaces (APIs). The platform would build on, rather than replicate existing integration and coordination efforts within the space situational awareness ecosystem, using existing standards for data message formats from organizations like the Consultative Committee for Space Data Systems and wrapping, rather than replacing existing integrations. We will present an initial STM architecture in this presentation, with a few examples showing how stakeholders can interact structurally, but flexibly, within this architecture.

Space Situational Awareness↗

Beyond Fair: Engagement, Data Usability, and Open Community Productivity through the NASA Open Science Data Repository

The FAIR principle (findable, accessible, interoperable, and reusable) governs the storage and sharing of NASA space biology and health data[1]. These guiding principles maximize reuse of data and the reproducibility of scientific findings. The NASA Open Science Data Repository (OSDR; an expansion of NASA GeneLab) was built on the FAIR principles and houses over 500 studies and close to 1000 datasets from decades of space life sciences experiments. OSDR embodies the FAIR principles through data governance that includes mediated, embargoed, and fully open access data. The FAIR data governance principles were recently proposed to be expanded to encompass a FAIREST framework for assessing research data repositories (FAIR + Engagement, Social connections, and Trust)[2]. FAIREST emphasizes the importance of data repositories engaging with the scientific community and gaining the trust of researchers regarding data quality. Trust also refers to the TRUST principles developed for assessment of digital repositories: Transparency, Responsibility, User Focus, Sustainability, Technology[3]. We present the “Open Science for Life in Space” Analysis Working Groups (AWGs) as evidence regarding the power of engagement, social connections, and trust which has enhanced OSDR’s capabilities and productivity. AWG members engage in two main activities. One, members provide feedback on OSDR scientific standards for data ingestion, curation, and reuse (study, subject and assay metadata; processing pipelines; dataset formats and uniformed structures for machine-readability). Two, AWG members collaborate to mine-reuse OSDR data to conduct scientific analysis. With nearly 800 active members, the AWGs have resulted in 32 publications re-using OSDR data and contributed many papers in two major special issues in Cell (2020) and Nature (2024). AWGs also serve as networking groups, facilitate social connections between researchers at all levels of experience, and also have a social online ‘Forum’ used to keep members informed on projects and opportunities. This community-centric, productive, and trustworthy data culture has resulted in a broader effect with international space agencies, academics, and the commercial space sector wanting to submit their data to OSDR. Ten studies of Inspiration 4 data were recently publicly released by OSDR, as were some JAXA human data. Coming up soon in OSDR are data submissions from the European Space Agency, Virgin Galactic PIs, and SpaceX Polaris Dawn. A major benefit of OSDR is the array of standardized and uniformly formatted data (which was developed through AWG member consensus), from which visualization tools, analysis tools, and machine learning models can be built or trained. This talk will cover the Multi-Study Visualization Tool, the Environmental Data Application, RadLab, and a UCSF-NSF funded knowledge graph biomedical health discovery tool ‘SPOKE’ currently being integrated with OSDR. OSDR also provides training programs in bioinformatics and machine learning to improve the scientific community’s awareness of data availability and to boost their ability to perform data analysis. The increasing engagement of the scientific community and the public with technologies powered by artificial intelligence (AI) heightens the need for data analysis to be transparent. The AI for Life in Space initiative leverages the data products provided in OSDR to train AI models, with an emphasis on explainable and trustworthy AI, which would not be possible without FAIR data and metadata. Overall, here we will demonstrate the importance for NASA life sciences data repositories to adhere to the FAIREST framework, by providing examples and success stories from different aspects of OSDR.

data↗

Intelligent Systems Technologies and Utilization of Earth Observation Data

The addition of raw data and derived geophysical parameters from several Earth observing satellites over the last decade to the data held by NASA data centers has created a data rich environment for the Earth science research and applications communities. The data products are being distributed to a large and diverse community of users. Due to advances in computational hardware, networks and communications, information management and software technologies, significant progress has been made in the last decade in archiving and providing data to users. However, to realize the full potential of the growing data archives, further progress is necessary in the transformation of data into information, and information into knowledge that can be used in particular applications. Sponsored by NASA s Intelligent Systems Project within the Computing, Information and Communication Technology (CICT) Program, a conceptual architecture study has been conducted to examine ideas to improve data utilization through the addition of intelligence into the archives in the context of an overall knowledge building system (KBS). Potential Intelligent Archive concepts include: 1) Mining archived data holdings to improve metadata to facilitate data access and usability; 2) Building intelligence about transformations on data, information, knowledge, and accompanying services; 3) Recognizing the value of results, indexing and formatting them for easy access; 4) Interacting as a cooperative node in a web of distributed systems to perform knowledge building; and 5) Being aware of other nodes in the KBS, participating in open systems interfaces and protocols for virtualization, and achieving collaborative interoperability.

Ramapriyan, H. K.↗

Opening Historical Airborne Data to Present Day Researchers

For more than 50 years, NASA has flown airborne sensors to carry out research, validate satellite sensors, and test new instrument capabilities. Data collected prior to 2000 are typically analog and difficult to locate and use. The Airborne Data Management Group (ADMG) facilitates rescue of these valuable data to ensure easier discovery, access, and use. But opening historical data comes at a cost of both time and money. Careful decisions are required in assessing the return on investment. - Is there interest in the science community? - Are there government data requirements? - What is the temporal / spatial value of the data? - Can data be transformed to a digital format? - What is cost of transformation? - What time period is needed for rescue? Converting the data to today’s digital storage standards increases value and provides data access. The addition of metadata makes the data easier to search for.

Deborah Smith↗

An Overview of NASA’s Catalog of Archived Suborbital Earth Science Investigations (CASEI): Supporting FAIR and Open Access to Airborne and Field Data

Since 2019, NASA’s Airborne Data Management Group (ADMG) within the Interagency Implementation and Advanced Concepts Team (IMPACT) has worked to promote and ensure the discoverability and accessibility of the agency’s non-satellite Earth science observations. A primary component of this effort is the development of NASA’s Catalog of Archived Suborbital Earth Science Investigations (CASEI) and the vetting of key contextual details required to sustain this unique inventory of airborne and field metadata. CASEI provides information on the science objectives motivating data collection, key events/time periods in the observational record aligned with the science objectives, complementary simultaneous observations, programmatic details, and much more. The diverse set of data formats and disciplines served by CASEI have required the implementation of a common data model to organize suborbital observation metadata and efficiently connect appropriate campaigns, platforms, and instruments. The CASEI inventory provides a single entry point for users to search and browse NASA’s airborne and field data archives, regardless of which repository is responsible for their stewardship. This presentation will provide a summary of the motivations for and the development of the CASEI system. Particular attention will be granted to how CASEI facilitates discovery and reuse of these lesser-known NASA data, supporting the Open Science vision and enhancing the return on investments made to collect these unique and varied observations. An up-to-date summary of CASEI inventory content and initial metrics will be provided. Current and future avenues ADMG is pursuing to enhance both CASEI and specific components of suborbital data stewardship at various stages of the data life cycle will also be discussed.

Stephanie M. Wingo↗

Time-History Statistics of Soot Formation in A Model Gas Turbine Combustor

Soot formation is a complex dynamic and intermittent process determined by properties of the fuel, combustor design, and combustor operation. Although the major steps in soot formation (i.e., formation of precursors, inception, growth and evolution) are similar for a variety of carbonaceous fuels, applications, and operating conditions, it remains unclear when the temporal transition between these steps occurs. An engineering prediction tool coupled with computational fluid physics (CFD), therefore needs to accurately model all these complex steps. To develop such a model, we propose the time-history concept for understanding the time dependency of soot formation as a function of local properties (i.e., temperature, velocity, local fuel air ratio, etc.). We continue our previous work with modeling the DLR aero-combustor [1] with our updated in-house CFD code, Open National Combustion Code (OpenNCC), that now includes a Multiple Time-Scale Flamelet Progress Variable approach and a the semi-empirical two-equation soot model. We injected massless tracer particles upstream of the injector region of the combustor to collect time-history statistics of the solution variables. The correlations between the collected statistics with respect to the experimental soot volume fraction data showed that time-history effect of certain flow variables, including turbulent kinetic energy (TKE), and multiple species is indeed important for soot formation. We then conducted a time-history based correlation analysis to determine the key species and the concentration ranges critical for soot formation (C6H5-based nucleation, acetylene-based surface growth, and oxidation with OH and O2). Based on the time-history correlation coefficient (THCC) analysis, we propose possible modifications to improve the current two-equation model.

LES↗

AI Curation Methods for NASA Scientific Data

The NASA Open Science Data Repository (OSDR) serves as a central hub for sharing and accessing NASA's vast collection of scientific data, supporting researchers across diverse fields. To enhance the efficiency, accuracy, and accessibility of this data, we are leveraging advanced artificial intelligence (AI) techniques as part of the AI for Curation project. By integrating large language models (LLMs) into our data curation workflow, we aim to streamline the entire process—from data submission to user interaction. This initiative focuses on improving key areas, including data ingestion, curation, and user engagement with curated datasets, impacting multiple domains and a wide user base. First, we are developing tools that can automatically parse data in various formats, using LLMs to convert unstructured data into structured, standardized formats. This reduces the manual effort required for curation, allowing curators to focus on more critical scientific analyses. Additionally, AI and machine learning (ML) models are being implemented to automate data validation and verification, ensuring the highest standards of data quality and reliability. Finally, we are creating a conversational AI agent to interact with the curated scientific studies in OSDR, helping users easily navigate the repository and access relevant data. By enhancing data discoverability and accessibility, these advancements will foster new research opportunities and promote the principles of open science.

Walter Alvarado↗

2022 Spring Internship Exit Presentation

As efforts of the National Aeronautics and Space Administration (NASA) and the Federal Aviation Administration (FAA) continue to digitize the air traffic management (ATM) domain, there is countless times of need for downstream natural language processing (NLP) tasks such as named entity recognition, text summarization, classification, and more. Although there are a plethora of open-sourced pre-trained transformer models in the NLP field such as BERT, RoBERTa, XLNet, and GPT-3, these models are trained on general corpora and perform poorly on domain-specific terminology and phraseology seen in ATM documents such as Notice to Airmen (NOTAMs) and Letters of Agreement (LoA). Our proposed research objective will be to first gather a large corpus of air traffic management related documents, orders, notices, books, technical papers, conference papers, articles, and other miscellaneous sources of text data from the FAA, NASA, and accredited conference and publication societies. After gathering this data, many steps will have to be taken to collate and preprocess the data into a format understandable by our test transformer models. Thirdly, we will set up training pipelines to train the RoBERTa model on its unsupervised training task masked language modelling (MLM) using resources provided by the NASA Advanced Supercomputing (NAS) facilities. Finally, these fine-tuned transformer models will be evaluated on their performance on down-stream NLP tasks as mentioned above, to show whether they will be effective when working with ATM related data or not. Once complete, this model could be made open-sourced on the HuggingFace website, where the rest of the ATM community can access and utilize this tool.

NLP↗

Doppler-corrected differential detection system

Doppler in a communication system operating with a multiple differential phase-shift-keyed format (MDPSK) creates an adverse phase shift in an incoming signal. An open loop frequency estimation is derived from a Doppler-contaminated incoming signal. Based upon the recognition that, whereas the change in phase of the received signal over a full symbol contains both the differentially encoded data and the Doppler induced phase shift, the same change in phase over half a symbol (within a given symbol interval) contains only the Doppler induced phase shift, and the Doppler effect can be estimated and removed from the incoming signal. Doppler correction occurs prior to the receiver's final output of decoded data. A multiphase system can operate with two samplings per symbol interval at no penalty in signal-to-noise ratio provided that an ideal low pass pre-detection filter is employed, and two samples, at 1/4 and 3/4 of the symbol interval T sub s, are taken and summed together prior to incoming signal data detection.

Simon, Marvin K.↗