Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data processing and logging”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Observations of small-scale turbulence in the atmosphere of Venus by Mariner 5

Information regarding small-scale turbulence in the Venus atmosphere is important and desirable because it contributes to understanding of the atmosphere's circulation. It is demonstrated that the radio occultation data of a flyby spacecraft such as Mariner 5 can provide valuable information on turbulence in the Venus atmosphere. Unlike previous studies of the Mariner 5 data, this paper is based on the frequency spectrum rather than the variance of the log-amplitude fluctuations. The excellent agreement between the processed and previously derived theoretical spectra furnishes strong evidence that the Mariner 5 fluctuations are primarily turbulence-induced. It is seen that, above 35 km, turbulence is strongest in the vicinity of 45 and 60 km, and that the outer scale of turbulence is of the order of 100 m. Comparison with the results obtained from the Venera missions is also discussed.

Woo, R.↗

Program Facilitates Distributed Computing

KNET computer program facilitates distribution of computing between UNIX-compatible local host computer and remote host computer, which may or may not be UNIX-compatible. Capable of automatic remote log-in. User communicates interactively with remote host computer. Data output from remote host computer directed to local screen, to local file, and/or to local process. Conversely, data input from keyboard, local file, or local process directed to remote host computer. Written in ANSI standard C language.

Hui, Joseph↗

Psychophysiological Assessment of Fatigue in Commercial Aviation Operations

The overall goal of this study is to improve our understanding of crew work hours, workload, sleep, fatigue, and performance, and the relationships between these variables on actual flight deck performance. Specifically, this study will provide objective measures of physiology and performance, which may benefit investigators in identifying fatigue levels of operators in commercial aviation and provide a way to better design strategies to limit crew fatigue. This research was supported by an agreement between NASA Ames Research Center and easyJet Airline Company, Ltd., Luton, UK. Twenty commercial pilots volunteered to participant in the study that included 15 flight duty days. Participants wore a Zephyr Bioharness ambulatory physiological monitor each flight day, which measured their heart rate, respiration rate, skin temperature, activity and posture. In addition, pilots completed sleep log diaries, self-report scales of mood, sleepiness and workload, and a Performance Vigilance Task (PVT). All data were sent to NASA researchers for processing and analyses. Heart rate variability data of several subjects were subjected to a spectral analysis to examine power in specific frequency bands. Increased power in low frequency band was associated with reports of higher subjective sleepinesss in some subjects. Analyses of other participants data are currently underway.

Hernandez, Norma↗

Restoration of the Apollo 15 Heat Flow Experiment Data from 1975 to 1977

The Apollo 15 Heat Flow Experiment (HFE) was conducted from July 1971 through January 1977. Two heat flow probes were deployed roughly 8.5 meters apart. Probe 1 and Probe 2 penetrated to 1.4-meters and 1-meter depths into the lunar regolith, respectively. Temperatures at different depths and the surface were logged with 7.25-minute intervals and transmitted to Earth. At the conclusion of the experiment, only data obtained from July 1971 through December 1974 were processed and archived at the National Space Science Data Center (NSSDC) by the principal investigator of the experiment, Marcus Langseth of Columbia University. Langseth died in 1997. It is not known what happened to the HFE data tapes he used. Current researchers have strong interests in re-examining the HFE data for the full duration of the experiment. We have recovered and processed large portions of the Apollo 15 HFE data from 1975 through 1977 by assembling data and metadata from various sources.

HFE↗

Exploiting the Free Landsat Archive for Operational Monitoring of Ecosystem Condition and Change Across the Chesapeake Bay Watershed

For the first time, all imagery acquired by the Landsat series of satellites is being made available by the USGS to users at no cost. This represents a key opportunity to use Landsat in a truly operational monitoring framework: large regions of the U.S. such as the Chesapeake Bay Watershed can now be analyzed using "wall-to-wall" imagery at timescales from approximately 1 month to several years. With the future launch of the Landsat Data Continuity Mission (LDCM) and Decadal Survey missions such as the hyperspectral HyspIRI, it is imperative to develop robust processing systems to perform annual ecosystem assessments over large regions such as the Chesapeake Bay. We have been working at NASA's Goddard Space Flight Center (GSFC) to develop an integrative framework for inserting 30m, annual, Landsat based data and derived products into the existing decision support system for the Bay, with a particular focus on ecosystem condition and changes over the entire watershed. The basic goal is to use a 'stack' of Landsat imagery with 40% or less cloud cover to produce multi-date (2005-2009 period), cloud/shadow/gap-free composited surface reflectance products that will support the creation of watershed scale land cover/ use products and the monitoring of ecosystem change across the Bay. Our scientific focus extends beyond the conventional definition of land cover (i.e. a classification of vegetation type) as we propose to monitor both changes in surface type (e.g. forest to urban), vegetation structure (e.g. forest disturbance due to logging or insect damage), as well as winter crop cover. These processes represent a continuum from large, interannual changes in land cover type, to subtler, intra-annual changes associated with short-term disturbance. The free Landsat data are being processed to surface reflectance and composited using the existing Landsat Ecosystem Disturbance Adaptive Processing System here at NASA/ GSFC, and land cover products (type, tree cover, impervious cover, winter cover) are being produced using well-established decision tree and regression tree algorithms. The goal of this session is to present the data products that we have been developing to the Bay science community and to discuss potential avenues for improvements and usage of the products for decision support.

BrowndeColstoun, Eric↗

Roughrider Carbon Storage Hub (Final Report)

The Roughrider Carbon Storage Hub was a 2-year project (October 2023 – September 2025) conducted by the Energy & Environmental Research Center (EERC) focused on advancing the feasibility of a commercial-scale carbon dioxide (CO 2 ) geologic storage hub in McKenzie County, North Dakota. The project’s objective was to investigate the potential that stacked storage complexes (multiple deep saline formations) can safely and economically store at least 50 million tonnes of CO 2 within 30 years. The captured CO 2 would be sourced from industrial emitters including project partner ONEOK, Inc.’s gas-processing plants and a planned gas-to-liquids facility. Drilling of the Roughrider 1 stratigraphic test well (14,979-ft total depth) was completed in November 2024. The wellbore intersected four candidate storage formations: Inyan Kara, Broom Creek, Mission Canyon, and Black Island–Deadwood. Operational challenges, including a stuck drill string, were resolved without long-term impact. A comprehensive logging and coring program was conducted, followed by successful well abandonment and site reclamation. Over 660 ft of 4-in. whole core was retrieved. Core plug samples were processed and analyzed for petrophysical and geochemical properties. Results confirmed promising porosity and permeability in the Inyan Kara and Broom Creek Formations and removal of the Mission Canyon and Black Island–Deadwood horizons from further investigation. Data derived from the logging and coring program were used to improve initial geologic models built from legacy data. CO 2 injection simulations showed that the Inyan Kara alone can feasibly store the target mass of CO 2 . Because of subtle differences in geologic structure and porosity trends between the formations, a stacked storage scenario using the Broom Creek and Inyan Kara Formations resulted in a larger overall plume area than using the Inyan Kara alone. Preliminary CO 2 pipeline routes from the industrial sources were mapped utilizing existing rights of way and evaluated for capacity and cost using U.S. Department of Energy Office of Fossil Energy and Carbon Management/National Energy Technology Laboratory models and U.S. Environmental Protection Agency emissions data. Integrating capture, transport, and storage cost estimates with policy incentives (e.g., 45Q credits) provided a total cost-per-ton analysis. Results indicate that the small scale of the volumes to be transported over the cumulative large distances does not support the project’s financial viability. However, the groundwork laid during this project from geological, regulatory, and social perspectives positions the Roughrider hub site as a promising candidate for commercial carbon storage in North Dakota, especially if the economy of scale is introduced for CO 2 transportation to the hub site.

01 COAL, LIGNITE, AND PEAT↗

Oscilloscope Data Push Program

This paper details the development of a Python program designed to automate the data acquisition and conversion for an oscilloscope for the purposes of a one-off/temporary data acquisition system for users that readily need data, and do not have the option of obtaining a Data Acquisition (DAQ) solution. Creating DAQ systems for analyzing a system requires expensive electronics and a dedicated team of engineers for support. Traditionally, manual data collection and processing are time consuming and prone to error. By automating these processes, the cost, efficiency and accuracy of data handling are improved upon. This project involves the creation of a program that interacts with the oscilloscope. During this interaction, there are various functions being performed such as the acquisition of waveform data via floating points, generating plots with the acquired wave points, and storing of floating points in a CSV file format for future reference and plotting purposes. While the initial aim of the project included continuous logging to a cloud database, this was deferred due to time constraints. The results portrayed an almost-instant rate of data collection with a buffer time, showcasing the potential for further integration and real-time data processing.

Osei-Tutu, Jason↗

The Flux of Carbon from Selective Logging, Fire, and Regrowth in Amazonia

The major goal of this work was to develop a spatial, process-based model (CARLUC) that would calculate sources and sinks of carbon from changes in land use, including logging and fire. The work also included Landsat data, together with fieldwork, to investigate fire and logging in three different forest types within Brazilian Amazonia. Results from these three activities (modeling, fieldwork, and remote sensing) are described, individually, below. The work and some of the personnel overlapped with research carried out by Dr. Daniel Nepstad's LBA team, and thus some of the findings are also reported in his summaries.

Houghton, R. A.↗

ASDC’s Python-Based Metadata Extraction Pipeline for Suborbital Campaigns

The FAIRness of data products, especially findability and accessibility depend on rich metadata which, when extracted, can allow for proper curation. Over the past few years, the Atmospheric Science Data Center (ASDC) suborbital science support team has developed a metadata extraction pipeline to ensure the required metadata can be retrieved systematically, effectively, and efficiently to ensure the data can be used by a broad community. The development of a pipeline has presented many, but necessary, challenges to support archival and distribution of ASDC’s 30+ suborbital missions. Though sufficient metadata is provided by instrument scientists, the metadata may not be readily machine actionable due to different formats and templates. Further complicating metadata extraction, our team has found that the nature of metadata can be quite diverse given the difference in measurement types, instruments, and measurement platforms. A metadata extraction pipeline has been developed to provide an efficient, plugin-in based, method for adding new parsers, a configuration system that lets non-developers customize how files are processed, and a system for identifying and logging metadata quality issues to ensure they are readily found and addressed. The metadata extraction pipeline identifies critical pieces of metadata that are needed to promote data FAIRness, including location, file revision, measurement start/end datetime and can be easily modified to extract further information (such as variables). Given the wide-ranging datasets, the pipeline has been modified to accommodate multiple file formats, including multiple versions of ICARTT (International Consortium for Atmospheric Research on Transport and Transformation), HDF (Hierarchical Data Format), netCDF (network Common Data Form), and multiple versions of the Ames File Format. The pipeline also supports building metadata for file formats that cannot have metadata easily extracted from them, such as PDF (Portable Document Format) and GIF (Graphics Interchange Format). The pipeline has allowed our team to maintain a consistent flow of data and metadata to archival and distribution services, ensuring the ASDC meets the needs of the suborbital science community. This presentation will highlight the ASDC’s suborbital metadata extraction pipeline, its development, how it’s been modified to support data FAIRness, and plans for maintaining the pipeline and adding new features.

Abraham Porter↗

Access Patterns and Performance Behaviors of Multi-layer Supercomputer I/O Subsystems under Production Load

Scientific computing workloads at HPC facilities have been shifting from traditional numerical simulations to AI/ML applications for training and inference while processing and producing ever-increasing amounts of scientific data. To address the growing need for increased storage capacity, lower access latency, and higher bandwidth, emerging technologies such as non-volatile memory are integrated into supercomputer I/O subsystems. With these emerging trends, we need a better understanding of the multilayer supercomputer I/O systems and ways to use these subsystems efficiently. In this work, we study the I/O access patterns and performance characteristics of two representative supercomputer I/O subsystems. Through an extensive analysis of year-long I/O logs on each system, we report new observations in I/O reads and writes, unbalanced use of storage system layers, and new trends in user behaviors at the HPC I/O middleware stack.

Bez, JL↗

Pseudonymization at Scale: OLCF’s Summit Usage Data Case Study

The analysis of vast amounts of data and the processing of complex computational jobs have traditionally relied upon high performance computing (HPC) systems, which offer reliable and efficient management of large-scale computational and data resources. Understanding these analyses’ needs is paramount for designing solutions that can lead to better science, and similarly, understanding the characteristics of the user behavior on those systems is important for improving user experiences on HPC systems. A common approach to gathering data about user behavior is to extract workload characteristics from system log data available only to system administrators. Recently at Oak Ridge Leadership Computing Facility (OLCF), however, we unveiled user behavior about the Summit supercomputer by collecting data from a user’s point of view with ordinary Unix commands.In this paper, we discuss the process, challenges, and lessons learned while preparing this dataset for publication and submission to an open data challenge. The original dataset contains personal identifiable information (PII) about the users of OLCF which needed be masked prior to publication, and we determined that anonymization, which scrubs PII completely, destroyed too much of the structure of the data to be interesting for the data challenge. We instead chose to pseudonymize the dataset, which reduced the linkability of the dataset to the users’ identities. Pseudonymization is significantly more computationally expensive than anonymization, and the size of our dataset, which is approximately 175 million lines of raw text, necessitated the development of a parallelized workflow that could be reused on different HPC machines. We demonstrate the scaling behavior of the workflow on two leadership class HPC systems at OLCF, and we show that we were able to bring the overall makespan time from an impractical 20+ hours on a single node down to around 2 hours. As a result of this work, we release the entire pseudonymized dataset and make the workflows and source code publicly available.

Maheshwari, Ketan↗

Estimating residual fault hitting rates by recapture sampling

For the recapture debugging design introduced by Nayak (1988) the problem of estimating the hitting rates of the faults remaining in the system is considered. In the context of a conditional likelihood, moment estimators are derived and are shown to be asymptotically normal and fully efficient. Fixed sample properties of the moment estimators are compared, through simulation, with those of the conditional maximum likelihood estimators. Properties of the conditional model are investigated such as the asymptotic distribution of linear functions of the fault hitting frequencies and a representation of the full data vector in terms of a sequence of independent random vectors. It is assumed that the residual hitting rates follow a log linear rate model and that the testing process is truncated when the gaps between the detection of new errors exceed a fixed amount of time.

Lee, Larry↗

pvOps: a Python package for empirical analysis of photovoltaic field data

The purpose of pvOps is to support empirical evaluations of data collected in the field related to the operations and maintenance (O&M) of photovoltaic (PV) power plants. pvOps presently contains modules that address the diversity of field data, including text-based maintenance logs, current-voltage (IV) curves, and timeseries of production information. The package functions leverage machine learning, visualization, and other techniques to enable cleaning, processing, and fusion of these datasets. These capabilities are intended to facilitate easier evaluation of field patterns and extraction of relevant insights to support reliability-related decision-making for PV sites. The open-source code, examples, and instructions for installing the package through PyPI can be accessed through the GitHub repository.

14 SOLAR ENERGY↗

Supplier Management System

Supplier Management System (SMS) allows for a consistent, agency-wide performance rating system for suppliers used by NASA. This version (2.0) combines separate databases into one central database that allows for the sharing of supplier data. Information extracted from the NBS/Oracle database can be used to generate ratings. Also, supplier ratings can now be generated in the areas of cost, product quality, delivery, and audit data. Supplier data can be charted based on real-time user input. Based on these individual ratings, an overall rating can be generated. Data that normally would be stored in multiple databases, each requiring its own log-in, is now readily available and easily accessible with only one log-in required. Additionally, the database can accommodate the storage and display of quality-related data that can be analyzed and used in the supplier procurement decision-making process. Moreover, the software allows for a Closed-Loop System (supplier feedback), as well as the capability to communicate with other federal agencies.

Ramirez, Eric↗

Mission Operations Center (MOC) - Precipitation Processing System (PPS) Interface Software System (MPISS)

MPISS is an automatic file transfer system that implements a combination of standard and mission-unique transfer protocols required by the Global Precipitation Measurement Mission (GPM) Precipitation Processing System (PPS) to control the flow of data between the MOC and the PPS. The primary features of MPISS are file transfers (both with and without PPS specific protocols), logging of file transfer and system events to local files and a standard messaging bus, short term storage of data files to facilitate retransmissions, and generation of file transfer accounting reports. The system includes a graphical user interface (GUI) to control the system, allow manual operations, and to display events in real time. The PPS specific protocols are an enhanced version of those that were developed for the Tropical Rainfall Measuring Mission (TRMM). All file transfers between the MOC and the PPS use the SSH File Transfer Protocol (SFTP). For reports and data files generated within the MOC, no additional protocols are used when transferring files to the PPS. For observatory data files, an additional handshaking protocol of data notices and data receipts is used. MPISS generates and sends to the PPS data notices containing data start and stop times along with a checksum for the file for each observatory data file transmitted. MPISS retrieves the PPS generated data receipts that indicate the success or failure of the PPS to ingest the data file and/or notice. MPISS retransmits the appropriate files as indicated in the receipt when required. MPISS also automatically retrieves files from the PPS. The unique feature of this software is the use of both standard and PPS specific protocols in parallel. The advantage of this capability is that it supports users that require the PPS protocol as well as those that do not require it. The system is highly configurable to accommodate the needs of future users.

Ferrara, Jeffrey↗

University participation via UNIDATA, part 2

The University Corporation for Atmospheric Research (UCAR) is presently completing UNIDATA, Phase II, considered to be the design phase of the UNIDATA Project. The four major components of the UNIDATA System are: (1) global services which access is provided, (2) long haul communication for providing that access, (3) local services for providing access and local management of acquired data, and (4) local interactive processing and graphical display. Each component is described in detail with linkages among the components elucidated. Within this framework, access to the PCDS is discussed. It is pointed out that access to the PCDS could occur via general purpose computer-to-computer communications providing remote log on to the system. The UNIDATA System could also be used to transfer information from the PCDS, provided the appropriate software is available to receive the data. Both of these scenarios require agreements on the access protocols and appropriate physical connections. Universities' needs for weather information on a near real-time basis, and UNIDATA has already established a satellite broadcast data service for this purpose.

Fulker, D. W.↗

G-LiHT Campaign Leaf Carbon and Nitrogen Content, Mar2017: Puerto Rico

Measurements of leaf carbon and nitrogen content collected from 68 tropical tree species. Data includes leaves collected from fully sunlit and shaded canopy strata as well as leaves for young, mature, old and senescent leaf ages. Data for each sample includes the relative age estimate, leaf canopy position and sample number. This data was collected as part of the 2017 NGEE-Tropics / NASA G-LiHT airborne campaign. This data package includes processed data for leaf carbon and nitrogen content (*.csv). Metadata files include data description (_dd.csv) for tabular data, site information (*.csv), sampling protocol (*.pdf) and the NGEE-Tropics FRAMES e-field log and file submission metadata (*.xlsx). See related datasets for sample details including photographs, leaf-level reflectance and transmittance spectra, leaf mass per area (LMA) and water content.

54 ENVIRONMENTAL SCIENCES↗

Visualizing Multi-process CPU Utilization using CUSP

The CPU Utilization Statistics Plotter (CUSP) tool automates the interpretation of detailed CPU Utilization trace data and statistics. It puts you on the cusp of understanding how CPU resources are split among the many parallel components of a software system.CUSP combines time-sampled CPU utilization numbers and Event Log annotations to generate human-readable plots and tables. It automatically splits up large CPU usage log files around interesting events, determines and highlights just the tasks of primary relevance by evaluating their changing contribution to each plot's total CPU usage, automatically eliminates irrelevant tasks, provides context by labeling plots with names and durations of all active commands, and uses consistent color-coding to enable quick visual comparison across multiple plots.CUSP has been used to process CPU Utilization trace logs on the Mars Science Laboratory and the Mars 2020 Rover missions during flight software development and Flight Operations on the Martian surface since December 2013.

Maimone, Mark W↗