Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “open data format”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Alpha-decay width of a near proton-threshold resonance in the 7 Li( α, α ) channel

We investigate the alpha-decay width of a near proton-threshold resonance in 11 B by the excitation functions of 7 Li(α, α) and 7 Li(α, α’) reactions. This resonance is an example of loosely bound atomic nuclei, understood as small open quantum systems, that exhibit properties significantly influenced by their coupling to the continuum. The present experiment focuses on the formation of a specific state in 11 B at excitation energy in the region near 11.4 MeV by bombarding a 7 Li target with a 4 He beam at energies ranging from 3.92 to 4.56 MeV in the laboratory frame. The study provides data complementary to previous observations of the proton emission in the β-decay of the neutron-rich halo nucleus 11 Be. An R-matrix fit to the data provides resonance energy and partial widths consistent with J π =1/2 + for a narrow near-threshold state in 11 B at 11.400(20) MeV, that is 171(20) keV above the proton emission threshold, with partial α-decay width Γ α =2.5$^{+3.5}_{–1.5}$ keV in elastic scattering and Γ α' =15.8$^{+1.9}_{–0.4}$ keV in inelastic scattering. These findings contribute to the understanding of the structure of this open quantum system resonance.

11B↗

Improving the Accessibility and Usability of Geothermal Information with Data Lakes and Data Pipelines on the Geothermal Data Repository: Preprint

The Geothermal Data Repository (GDR) provides universal access to data and information resulting from research and development activities funded by the Department of Energy (DOE). The GDR has extended this universal access to big data through integration with data lakes developed by the Open Energy Data Initiative (OEDI). Previously, large datasets such as seismic waveform or distributed acoustic sensing (DAS) data could only be accessed by institutions with high performance data storage and compute capabilities, effectively limiting the accessibility of big data to national labs, larger universities, and major corporations. Moreover, the time and resources needed to transport big data and configure them can produce additional barriers to use. Many of the standard formats used for structured data models (also known as content models) are incapable of handling big data and can introduce additional usability problems, often requiring data to be reformatted prior to use. This paper will explore how recent integrations between the GDR and the OEDI data lake have improved the accessibility and usability of geothermal data in a big way, making the data available to a broader audience, and enabling collaborative analysis and innovation across the greater geothermal industry.

access↗

Improving the Accessibility and Usability of Geothermal Information with Data Lakes and Data Pipelines on the Geothermal Data Repository

The Geothermal Data Repository (GDR) provides universal access to data and information resulting from research and development activities funded by the Department of Energy (DOE). The GDR has extended this universal access to big data through integration with data lakes developed by the Open Energy Data Initiative (OEDI). Previously, large datasets such as seismic waveform or distributed acoustic sensing (DAS) data could only be accessed by institutions with high performance data storage and compute capabilities, effectively limiting the accessibility of big data to national labs, larger universities, and major corporations. Moreover, the time and resources needed to transport big data and configure them can produce additional barriers to use. Many of the standard formats used for structured data models (also known as content models) are incapable of handling big data and can introduce additional usability problems, often requiring data to be reformatted prior to use. This paper will explore how recent integrations between the GDR and the OEDI data lake have improved the accessibility and usability of geothermal data in a big way, making the data available to a broader audience, and enabling collaborative analysis and innovation across the greater geothermal industry.

access↗

PlasmoData.jl — A Julia framework for modeling and analyzing complex data as graphs

Datasets encountered in scientific and engineering applications appear in complex formats (e.g., images, multivariate time series, molecules, video, text strings, networks). Graph theory provides a unifying framework to model such datasets and enables the use of powerful tools that can help analyze, visualize, and extract value from data. In this work, we present PlasmoData.jl, an open-source, Julia framework that uses concepts of graph theory to facilitate the modeling and analysis of complex datasets. The core of our framework is a general data modeling abstraction, which we call a DataGraph. We show how the abstraction and software implementation can be used to represent diverse data objects as graphs and to enable the use of tools from topology, graph theory, and machine learning (e.g., graph neural networks) to conduct a variety of tasks. We illustrate the versatility of the framework by using real datasets: (i) an image classification problem using topological data analysis to extract features from the graph model to train machine learning models; (ii) a disease outbreak problem where we model multivariate time series as graphs to detect abnormal events; and (iii) a technology pathway analysis problem where we highlight how we can use graphs to navigate connectivity. Further, our discussion also highlights how PlasmoData.jl leverages native Julia capabilities to enable compact syntax, scalable computations, and interfaces with diverse packages. Overall, we show that the DataGraph abstraction and PlasmoData.jl Julia package are able to model data within graphs and enable useful analysis.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Astronomical Data Center Bulletin, volume 1, number 2

Work in progress on astronomical catalogs is presented in 16 papers. Topics cover astronomical data center operations; automatic astronomical data retrieval at GSFC; interactive computer reference search of astronomical literature 1950-1976; formatting, checking, and documenting machine-readable catalogs; interactive catalog of UV, optical, and HI data for 201 Virgo cluster galaxies; machine-readable version of the general catalog of variable stars, third edition; galactic latitude and magnitude distribution of two astronomical catalogs; the catalog of open star clusters; infrared astronomical data base and catalog of infrared observations; the Air Force geophysics laboratory; revised magnetic tape of the N30 catalog of 5,268 standard stars; positional correlation of the two-micron sky survey and Smithsonian Astrophysical Observatory catalog sources; search capabilities for the catalog of stellar identifications (CSI) 1979 version; CSI statistics: blue magnitude versus spectral type; catalogs available from the Astronomical Data Center; and status report on machine-readable astronomical catalogs.

Nagy, T. A.↗

Optimum Detection of Frequency-Hopped Signals

This paper derives and analyzes optimum and near-optimum structures for detecting frequency-hopped (FH) signals with arbitrary modulation in additive white Gaussian noise. The principalmodulation formats considered are M-ary frequency-shift-keying (MFSK) with fast frequency hopping(FFH) wherein a single tone is transmitted per hop, and slow frequency hopping (SFH) with multipleMFSK tones (data symbols) per hop. The SFH detection category has not previously been addressedin the open literature and its analysis is generally more complex than FFH.

Cheng, Unjeng↗

A Concept for Civil Space Traffic Management

As technology has improved, operators have sought to use cubesats, as well as smallsats more generally, to perform increasingly more ambitious and sophisticated functions. Despite this, practical concerns associated with cubesat infant mortality, conjunctions, limited maneuverability, and debris generation have been relatively muted because most cubesats have been launched to lower orbits that limit both their orbital lifetime and consequences should a collision occur. NASA ARC has developed a concept for a highly-automated and distributed space traffic management (STM) architecture, drawing on similar work done to provide traffic management for small unmanned aerial systems (UAS) operating at low altitudes. The system proposes a strategy to accommodate growing space traffic volume safely, as well as pave the way for a transition of civil STM authority to a civilian governmental entity. The architecture envisions an open-access software platform architecture of data and service suppliers, consumers, and regulators, connected via a set of application programming interfaces (APIs). The platform would build on, rather than replicate existing integration and coordination efforts within the space situational awareness ecosystem, using existing standards for data message formats from organizations like the Consultative Committee for Space Data Systems and wrapping, rather than replacing existing integrations. We will present an initial STM architecture in this presentation, with a few examples showing how stakeholders can interact structurally, but flexibly, within this architecture.

Space Situational Awareness↗

Beyond Fair: Engagement, Data Usability, and Open Community Productivity through the NASA Open Science Data Repository

The FAIR principle (findable, accessible, interoperable, and reusable) governs the storage and sharing of NASA space biology and health data[1]. These guiding principles maximize reuse of data and the reproducibility of scientific findings. The NASA Open Science Data Repository (OSDR; an expansion of NASA GeneLab) was built on the FAIR principles and houses over 500 studies and close to 1000 datasets from decades of space life sciences experiments. OSDR embodies the FAIR principles through data governance that includes mediated, embargoed, and fully open access data. The FAIR data governance principles were recently proposed to be expanded to encompass a FAIREST framework for assessing research data repositories (FAIR + Engagement, Social connections, and Trust)[2]. FAIREST emphasizes the importance of data repositories engaging with the scientific community and gaining the trust of researchers regarding data quality. Trust also refers to the TRUST principles developed for assessment of digital repositories: Transparency, Responsibility, User Focus, Sustainability, Technology[3]. We present the “Open Science for Life in Space” Analysis Working Groups (AWGs) as evidence regarding the power of engagement, social connections, and trust which has enhanced OSDR’s capabilities and productivity. AWG members engage in two main activities. One, members provide feedback on OSDR scientific standards for data ingestion, curation, and reuse (study, subject and assay metadata; processing pipelines; dataset formats and uniformed structures for machine-readability). Two, AWG members collaborate to mine-reuse OSDR data to conduct scientific analysis. With nearly 800 active members, the AWGs have resulted in 32 publications re-using OSDR data and contributed many papers in two major special issues in Cell (2020) and Nature (2024). AWGs also serve as networking groups, facilitate social connections between researchers at all levels of experience, and also have a social online ‘Forum’ used to keep members informed on projects and opportunities. This community-centric, productive, and trustworthy data culture has resulted in a broader effect with international space agencies, academics, and the commercial space sector wanting to submit their data to OSDR. Ten studies of Inspiration 4 data were recently publicly released by OSDR, as were some JAXA human data. Coming up soon in OSDR are data submissions from the European Space Agency, Virgin Galactic PIs, and SpaceX Polaris Dawn. A major benefit of OSDR is the array of standardized and uniformly formatted data (which was developed through AWG member consensus), from which visualization tools, analysis tools, and machine learning models can be built or trained. This talk will cover the Multi-Study Visualization Tool, the Environmental Data Application, RadLab, and a UCSF-NSF funded knowledge graph biomedical health discovery tool ‘SPOKE’ currently being integrated with OSDR. OSDR also provides training programs in bioinformatics and machine learning to improve the scientific community’s awareness of data availability and to boost their ability to perform data analysis. The increasing engagement of the scientific community and the public with technologies powered by artificial intelligence (AI) heightens the need for data analysis to be transparent. The AI for Life in Space initiative leverages the data products provided in OSDR to train AI models, with an emphasis on explainable and trustworthy AI, which would not be possible without FAIR data and metadata. Overall, here we will demonstrate the importance for NASA life sciences data repositories to adhere to the FAIREST framework, by providing examples and success stories from different aspects of OSDR.

data↗

MBARI-WEC September and October 2022 Field Data

This data is needed to simulate a model of the MBARI-WEC (Monterey Bay Aquarium Research Institute, Wave Energy Converter device) in a simulation environment (e.g. Gazebo) for 56 observation dates in the time between September and October 2022, and to compare the simulation outputs to the corresponding field data of the physical MBARI-WEC. There were 50 observations chosen in Sept and 6 observations in Oct. To help understand terms below, a summary of the system can be found at the github link in the downloads section below. The Gazebo MBARI-WEC model is also provided, should users wish to simulate using this platform. There are 4 *mat files included. ................................................................................................................................................................................................................................... Spectrum and Simulation Inputs: September2022_spectrum_siminputs.mat and October2022_spectrum_siminputs.mat has data needed for simulation inputs in table format. These include the ocean spectrum for an observation and operating parameters of the MBARI-WEC during that observation. They are organized as rows representing an observation and columns representing data. For example, for the September *mat there are 50 rows. The first 7 columns are Datetime, sig_waveheight, peak_period, mean_period, heaveconedoor_status, pistonpos_mean, and scale_factor: - Datetime is the date and time the observation occurred in PST - sig_waveheight is the significant wave height of the ocean spectrum during that observation in meters - peak_period is the peak period of the ocean spectrum during that observation in seconds - mean_period is the mean period of the ocean spectrum during that observation in seconds - heaveconedoor_status is the status of the heave cone doors where 0 represents the doors are open and 1 represents they are closed - pistonpos_mean is the mean position of the PTO ram (piston) in meters - scale_factor is an additional factor of 0.5 --1.4 applied to a default damping relationship The next columns are data needed to represent the ocean spectrum. First are the frequencies [Hz] labeled as "f0-f38", then the variance density [m2/Hz] labeled as "vardens0-vardens38". October2022_spectrum_siminputs.mat follows as a similar format as above, but includes a larger amount of ocean spectrum frequencies and variance density elements. ................................................................................................................................................................................................................................... Field data: The field data is found in MBARIWEC_septdata.mat and MBARIWEC_octdata.mat for the observations of September and October, respectively. These contain data in a struct format. The struct contains the following fields for each observation: PC_BattCurr, PC_LoadCurr, PC_RPM, PC_Voltage, SC_Range, SC_Velocity, DateTime, where: - PC_BattCurr is the current flowing to or from the onboard batteries in Amps - PC_LoadCurr is the current flowing to the load dump in Amps - PC_RPM is the electric/hydraulic motor shaft speed (directly coupled) in RPM - PC_Voltage is the bus voltage at the power converter in Volts - SC_Range is the PTO ram (piston) position in meters where 0 is fully retracted and 2.03 is fully extended - SC_Velocity is the PTO ram (piston) velocity in meters/sec - DateTime is the date and time of the sampled field data in each observation in PST - Electric Power is equal to: PC_Voltage*(PC_BattCurr + PC_LoadCurr) in Watts For example, upon loading MBARIWEC_octdata.mat, the aforementioned fields would be loaded, each with {6x1} cells for the 6 observations chosen in October. Within the first cell of e.g. SC_Range would be sampled data representing the field data of the MBARIWEC PTO piston position for say, one hour, of the first October observation. The corresponding field DateTime would...

16 TIDAL AND WAVE POWER↗

The "PVLib" of Degradation: PVDeg

The Photovoltaic (PV) industry constantly aims for lower costs through higher-efficiency cells, improved module designs, and improvements in durability. This leads to the use of new materials, designs, and manufacturing processes, and not always with a sufficient amount of durability testing. To help drive down costs there is a desire to create modules that will last for up to 50 years of service life. To accomplish this, every degradation mode and mechanism must be identified and either eliminated or otherwise mitigated. This involves the extrapolation of laboratory results to the field conditions. There is a need to organize the existing degradation data into an accessible format and to provide industry relevant tools for extrapolation from laboratory to field conditions. While the basic equations used to model degradation are sometimes very simple, the full analysis involves calculations are cumbersome but ubiquitous for many degradation processes. A simplified, modeling framework to accomplish these repetitive processes will facilitate the analysis to help researchers keep up with the rapid pace of technological changes. In this talk, we will describe our progress creating the open-source tool PVDeg. This tool can be used to search for and analyze degradation information and extrapolate PV module performance and durability to field exposure. PVDeg simplifies many of the common foundational computational operations for obtaining meteorological data and using it to generate a model of the PV deployment. This prediction tool repository also contains various degradation models as well as a library of material parameters suitable for estimating the durability assessment of materials and components. We use an integration pipeline approach that allows us to leverage weather data from the National Solar Radiation Database, and other weather sources, to perform geospatial degradation analysis in the US and worldwide. We hope to become a repository that can be used for weathering and degradation analysis for various applications beyond the PV industry. During the talk, we will provide the PVPMC attendees the opportunity to interact with the tool via a Google Collab tutorial they can run on their phones or laptops.

durability↗

Intelligent Systems Technologies and Utilization of Earth Observation Data

The addition of raw data and derived geophysical parameters from several Earth observing satellites over the last decade to the data held by NASA data centers has created a data rich environment for the Earth science research and applications communities. The data products are being distributed to a large and diverse community of users. Due to advances in computational hardware, networks and communications, information management and software technologies, significant progress has been made in the last decade in archiving and providing data to users. However, to realize the full potential of the growing data archives, further progress is necessary in the transformation of data into information, and information into knowledge that can be used in particular applications. Sponsored by NASA s Intelligent Systems Project within the Computing, Information and Communication Technology (CICT) Program, a conceptual architecture study has been conducted to examine ideas to improve data utilization through the addition of intelligence into the archives in the context of an overall knowledge building system (KBS). Potential Intelligent Archive concepts include: 1) Mining archived data holdings to improve metadata to facilitate data access and usability; 2) Building intelligence about transformations on data, information, knowledge, and accompanying services; 3) Recognizing the value of results, indexing and formatting them for easy access; 4) Interacting as a cooperative node in a web of distributed systems to perform knowledge building; and 5) Being aware of other nodes in the KBS, participating in open systems interfaces and protocols for virtualization, and achieving collaborative interoperability.

Ramapriyan, H. K.↗

Opening Historical Airborne Data to Present Day Researchers

For more than 50 years, NASA has flown airborne sensors to carry out research, validate satellite sensors, and test new instrument capabilities. Data collected prior to 2000 are typically analog and difficult to locate and use. The Airborne Data Management Group (ADMG) facilitates rescue of these valuable data to ensure easier discovery, access, and use. But opening historical data comes at a cost of both time and money. Careful decisions are required in assessing the return on investment. - Is there interest in the science community? - Are there government data requirements? - What is the temporal / spatial value of the data? - Can data be transformed to a digital format? - What is cost of transformation? - What time period is needed for rescue? Converting the data to today’s digital storage standards increases value and provides data access. The addition of metadata makes the data easier to search for.

Deborah Smith↗

An Overview of NASA’s Catalog of Archived Suborbital Earth Science Investigations (CASEI): Supporting FAIR and Open Access to Airborne and Field Data

Since 2019, NASA’s Airborne Data Management Group (ADMG) within the Interagency Implementation and Advanced Concepts Team (IMPACT) has worked to promote and ensure the discoverability and accessibility of the agency’s non-satellite Earth science observations. A primary component of this effort is the development of NASA’s Catalog of Archived Suborbital Earth Science Investigations (CASEI) and the vetting of key contextual details required to sustain this unique inventory of airborne and field metadata. CASEI provides information on the science objectives motivating data collection, key events/time periods in the observational record aligned with the science objectives, complementary simultaneous observations, programmatic details, and much more. The diverse set of data formats and disciplines served by CASEI have required the implementation of a common data model to organize suborbital observation metadata and efficiently connect appropriate campaigns, platforms, and instruments. The CASEI inventory provides a single entry point for users to search and browse NASA’s airborne and field data archives, regardless of which repository is responsible for their stewardship. This presentation will provide a summary of the motivations for and the development of the CASEI system. Particular attention will be granted to how CASEI facilitates discovery and reuse of these lesser-known NASA data, supporting the Open Science vision and enhancing the return on investments made to collect these unique and varied observations. An up-to-date summary of CASEI inventory content and initial metrics will be provided. Current and future avenues ADMG is pursuing to enhance both CASEI and specific components of suborbital data stewardship at various stages of the data life cycle will also be discussed.

Stephanie M. Wingo↗

Time-History Statistics of Soot Formation in A Model Gas Turbine Combustor

Soot formation is a complex dynamic and intermittent process determined by properties of the fuel, combustor design, and combustor operation. Although the major steps in soot formation (i.e., formation of precursors, inception, growth and evolution) are similar for a variety of carbonaceous fuels, applications, and operating conditions, it remains unclear when the temporal transition between these steps occurs. An engineering prediction tool coupled with computational fluid physics (CFD), therefore needs to accurately model all these complex steps. To develop such a model, we propose the time-history concept for understanding the time dependency of soot formation as a function of local properties (i.e., temperature, velocity, local fuel air ratio, etc.). We continue our previous work with modeling the DLR aero-combustor [1] with our updated in-house CFD code, Open National Combustion Code (OpenNCC), that now includes a Multiple Time-Scale Flamelet Progress Variable approach and a the semi-empirical two-equation soot model. We injected massless tracer particles upstream of the injector region of the combustor to collect time-history statistics of the solution variables. The correlations between the collected statistics with respect to the experimental soot volume fraction data showed that time-history effect of certain flow variables, including turbulent kinetic energy (TKE), and multiple species is indeed important for soot formation. We then conducted a time-history based correlation analysis to determine the key species and the concentration ranges critical for soot formation (C6H5-based nucleation, acetylene-based surface growth, and oxidation with OH and O2). Based on the time-history correlation coefficient (THCC) analysis, we propose possible modifications to improve the current two-equation model.

LES↗

AI Curation Methods for NASA Scientific Data

The NASA Open Science Data Repository (OSDR) serves as a central hub for sharing and accessing NASA's vast collection of scientific data, supporting researchers across diverse fields. To enhance the efficiency, accuracy, and accessibility of this data, we are leveraging advanced artificial intelligence (AI) techniques as part of the AI for Curation project. By integrating large language models (LLMs) into our data curation workflow, we aim to streamline the entire process—from data submission to user interaction. This initiative focuses on improving key areas, including data ingestion, curation, and user engagement with curated datasets, impacting multiple domains and a wide user base. First, we are developing tools that can automatically parse data in various formats, using LLMs to convert unstructured data into structured, standardized formats. This reduces the manual effort required for curation, allowing curators to focus on more critical scientific analyses. Additionally, AI and machine learning (ML) models are being implemented to automate data validation and verification, ensuring the highest standards of data quality and reliability. Finally, we are creating a conversational AI agent to interact with the curated scientific studies in OSDR, helping users easily navigate the repository and access relevant data. By enhancing data discoverability and accessibility, these advancements will foster new research opportunities and promote the principles of open science.

Walter Alvarado↗

2022 Spring Internship Exit Presentation

As efforts of the National Aeronautics and Space Administration (NASA) and the Federal Aviation Administration (FAA) continue to digitize the air traffic management (ATM) domain, there is countless times of need for downstream natural language processing (NLP) tasks such as named entity recognition, text summarization, classification, and more. Although there are a plethora of open-sourced pre-trained transformer models in the NLP field such as BERT, RoBERTa, XLNet, and GPT-3, these models are trained on general corpora and perform poorly on domain-specific terminology and phraseology seen in ATM documents such as Notice to Airmen (NOTAMs) and Letters of Agreement (LoA). Our proposed research objective will be to first gather a large corpus of air traffic management related documents, orders, notices, books, technical papers, conference papers, articles, and other miscellaneous sources of text data from the FAA, NASA, and accredited conference and publication societies. After gathering this data, many steps will have to be taken to collate and preprocess the data into a format understandable by our test transformer models. Thirdly, we will set up training pipelines to train the RoBERTa model on its unsupervised training task masked language modelling (MLM) using resources provided by the NASA Advanced Supercomputing (NAS) facilities. Finally, these fine-tuned transformer models will be evaluated on their performance on down-stream NLP tasks as mentioned above, to show whether they will be effective when working with ATM related data or not. Once complete, this model could be made open-sourced on the HuggingFace website, where the rest of the ATM community can access and utilize this tool.

NLP↗

Doppler-corrected differential detection system

Doppler in a communication system operating with a multiple differential phase-shift-keyed format (MDPSK) creates an adverse phase shift in an incoming signal. An open loop frequency estimation is derived from a Doppler-contaminated incoming signal. Based upon the recognition that, whereas the change in phase of the received signal over a full symbol contains both the differentially encoded data and the Doppler induced phase shift, the same change in phase over half a symbol (within a given symbol interval) contains only the Doppler induced phase shift, and the Doppler effect can be estimated and removed from the incoming signal. Doppler correction occurs prior to the receiver's final output of decoded data. A multiphase system can operate with two samplings per symbol interval at no penalty in signal-to-noise ratio provided that an ideal low pass pre-detection filter is employed, and two samples, at 1/4 and 3/4 of the symbol interval T sub s, are taken and summed together prior to incoming signal data detection.

Simon, Marvin K.↗

Research Report: Building a Wide Reach Corpus for Secure Parser Development

Whether developing from a specification or deriving parsers from samples, LangSec parser developers require widereach corpora of their target file format in order to identify key edge cases or common deviations from the format’s specification. In this work-in-progress paper, we report the details of several methods we’ve used to gather 30 million files, extract features and make these features amenable to search and other analytics. This paper documents opportunities and limitations of some popular open source data and tools and this paper will benefit researchers who need to efficiently gather a large file corpus.

Timmaraju, Virisha↗