Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data processing and logging”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

A Performant, Scalable Processing Pipeline for High‐Quality and FAIR Environmental Sensor Data

High-resolution environmental monitoring is necessary to record, understand, and predict biogeochemical and ecological changes particularly in coastal systems but brings significant challenges in processing and making rapidly available the resulting data. The COMPASS-FME project established a network of coastal observational sites across the Chesapeake Bay and western Lake Erie regions extensively instrumented with soil, vegetation, and weather sensors logging data every 15 min. Our data processing framework, written in R and completely open source, prioritizes rapid model-experiment iteration and makes biogeochemical data rapidly available for quality assurance/quality control, analysis, and model ingestion. This pipeline is distinguished by a standardized and modular approach to data curation, extensive metadata and documentation, and its high performance. These attributes combine to make biogeochemical data rapidly accessible across COMPASS-FME and the broader community. Flexible, powerful, and reproducible approaches to handling high-volume environmental data are crucial for accelerating biogeosciences research.

Pennington, Stephanie C. [Pacific Northwest Nation↗

April 2020 Darshan counters from the Summit supercomputer

This dataset is the Darshan counters collected from the Summit supercomputer in a month of April 2020. 1. Description of methods used for collection/generation of data: Job submitted on Summit HPC system when completed successfully and has made I/O calls (captured by Darshan tool) writes a Darshan log file on alpine filesystem. One job can have multiple `jsrun` commands and Darshan will generate separate logs each log corresponding to an `jsrun` command, so a job can have one or more Darshan logs associated with it. 2. Methods for processing the data: To process the data, we first use `darshan-util` tool to parse the Darshan logs. Then we restructure the logs and merge data from multiple Darshan logs if they belong to the same Summit job.

97 MATHEMATICS AND COMPUTING↗

Analytics-at-scale of Sensor Data for Digital Monitoring in Nuclear Plants (4th Annual Report)

Nuclear plant sites collect and store large volumes of data collected from various equipment and systems. These datasets typically include plant process parameters, maintenance records, technical logs, online monitoring data, and equipment failure data. The collection of such data affords an opportunity to leverage data-driven machine learning and artificial intelligence technologies to provide diagnostic and prognostic capabilities within the nuclear power industry to reduce operating and maintenance costs. In this way, nuclear energy can become more economically competitive with other energy sources, and premature closures can be avoided. From a maintenance standpoint, savings can be achieved by leveraging machine learning and artificial intelligence technologies to develop data-driven algorithms to better diagnose and predict potential faults within the system. Improved model accuracy can lead to reductions in unnecessary maintenance and more efficient planning of future maintenance, thus lowering the costs associated with parts, labor, and unnecessary planned, forced, or extended outages. From an operations perspective, cost savings can be generated by shifting from route-based monitoring to wireless technologies for online monitoring, and by transitioning from onsite- to cloud-based computing and storage services. Wireless monitoring would reduce the operator manhours required for taking routine measurements, while cloud computing services would generate cost savings by reducing the amount of hardware needing to be purchased and maintained—all while scaling to both computational and storage demands. This report summarizes this project’s effort to shift from costly, labor-intensive preventative maintenance to cheaper predictive maintenance.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Application of the informational reference system OZhUR to the automated processing of data from satellites of the Kosmos series

The structure and potential of the information reference system OZhUR designed for the automated data processing systems of scientific space vehicles (SV) is considered. The system OZhUR ensures control of the extraction phase of processing with respect to a concrete SV and the exchange of data between phases.The practical application of the system OZhUR is exemplified in the construction of a data processing system for satellites of the Cosmos series. As a result of automating the operations of exchange and control, the volume of manual preparation of data is significantly reduced, and there is no longer any need for individual logs which fix the status of data processing. The system Ozhur is included in the automated data processing system Nauka which is realized in language PL-1 in a binary one-address system one-state (BOS OS) electronic computer.

Pokras, V. M.↗

Assessment of Crew Time for Maintenance and Repairs Activities for Lunar Surface Missions

NASA is currently evaluating different methods to predict how much time crewmembers will spend conducting repair and maintenance activities on future space missions. As mission scope and spacecraft architectures change, it will be necessary to understand how crew repair and maintenance timelines are impacted by mission operations and technology changes. Past work has been done using historical ISS data to accurately predict crew habitation and operation timelines, resulting in the development of NASA’s Exploration Crew Time Model (ECTM). However, understanding crew maintenance and repair requirements has posed a unique challenge due to the complexity of available datasets, the probabilistic nature of sub-system failures, and the impacts of reliability growth on failure rates. This paper presents a methodology to collect and condition empirical repair and maintenance time data from available data sets, to extrapolate from that data to estimate projected maintenance and repair times for a lunar Surface Habitat, and to assess how uncertainty in repair time could impact utilization time on the lunar surface. NASA International Space Station (ISS) maintenance and crew time data are logged into two central databases, the Maintenance Data Collection (MDC) and the Operations Planning Timeline Integration System (OPTimIS) respectively. Separately, each of these two datasets capture only portions of the complete set of data required to generate an accurate assessment of crew time spent on maintenance activities at a sub-system level. MDC provides a detailed catalog of failure events and an overview of the failure’s required maintenance and OPTimIS provides a description of crew activities and crew time durations dedicated to maintenance. To create a more useful crew time estimate for maintenance timelines, the authors developed a methodology to capture relevant data from each set and combine and utilize that data by linking crew time requirements to specific components. The authors compare the failure logs in the MDC to crew activity logs pulled from OPTimIS and then process the data to estimate required repair times for each failure event. Data is also classified by the outcome of each repair event, whether the failed component was replaced or whether it was repaired in place. The entire maintenance activity dataset is then categorized based on the class of failed component to allow for a statistically significant sample size for each class and to provide accurate crew time estimates for any components lacking relevant data. This resultant component repair time data can be used in the future to generate Mean Time To Repair (MTTR) estimates and confidence intervals for each class of component based on a probabilistic distribution of documented maintenance events. These improved MTTR values can then be applied to candidate element sub-system architectures, along with component Mean Time Between Failure (MTBF) data to generate distributions for potential required system crew repair time estimates for a given mission. Repair time distributions can then be used to develop more accurate crew schedules and to assess potential available utilization time.

Crew Time↗

Assessment of Crew Time for Maintenance and Repair Activities for Lunar Surface Missions

NASA is currently evaluating different methods to predict how much time crewmembers will spend conducting repair and maintenance activities on future space missions. As mission scope and spacecraft architectures change, understanding how crew repair and maintenance timelines are impacted by mission operations and technology changes is vital for future mission planning. Past work has been done using historical International Space Station (ISS) data to accurately predict crew habitation and operation timelines, resulting in the development of NASA’s Exploration Crew Time Model (ECTM). However, understanding crew maintenance and repair requirements has posed a unique challenge due to the complexity of available datasets, the probabilistic nature of sub-system failures, and the impacts of reliability growth on failure rates. This paper presents a methodology to collect and condition empirical repair and maintenance time data from available datasets, to extrapolate from that data to estimate projected maintenance and repair times for a lunar Surface Habitat (SH), and to assess how uncertainty in repair time could impact utilization time on the lunar surface. NASA ISS maintenance and crew time data are logged into two central databases: the Maintenance Data Collection (MDC) and the Operations Planning Timeline Integration System (OPTimIS). Separately, each of these two datasets capture only portions of the complete set of data required to generate an accurate assessment of crew time spent on maintenance activities at a sub-system level. To create a more useful crew time estimate for maintenance timelines, the authors developed a methodology to capture relevant data from each set and combine and utilize that data by linking crew time requirements to specific components. The authors compare the failure logs in the MDC to crew activity logs pulled from OPTimIS and then process the data to estimate required repair time for each failure and repair event. The entire maintenance activity dataset is then categorized based on the class of failed component to ensure a significant sample size for each class and accurate crew time estimates for any components lacking relevant data. This resultant component repair time data can be used in the future to generate Mean Time to Repair (MTTR) estimates and confidence intervals for each class of component based on a probabilistic distribution of documented maintenance events. These improved MTTR values can then be applied to candidate element sub-system architectures, along with component Mean Time Between Failure (MTBF) data to generate distributions for potential required system crew repair time estimates for a given mission. The authors applied these modeling methods to a case study of a crewed mission to the planned SH and produced expected corrective maintenance crew time distributions. The results produced an expected corrective maintenance crew time at over 24 hours per mission, and a maintenance crew time distribution that reflects the importance of planning for sufficient maintenance requirements each mission. Repair time distributions can then be used to develop more accurate crew schedules and to assess potential available utilization time.

Crew Time↗

An experimental search for near-wall boundary conditions for large eddy simulation

Instantaneous wall shear stress and streamwise velocities have been measured simultaneously in a flat plate, turbulent boundary layer at moderate Reynolds number in an effort to provide experimental support for large eddy simulations. Data were obtained by using a buried-wire wall shear gage and a hot-wire rake positioned in the log region of the flow. All data processing was accomplished with digital data analysis techniques on a minicomputer. Fluctuations of the instantaneous U plus versus Y plus profiles about a mean law of the wall are shown to be significant and complex. Peak cross-correlation values between wall shear stress and the velocities are high and reflect the passage of a large structure inclined at a small angle to the wall. Estimates of this angle are consistent with those made by other investigators. Conditional sampling techniques were used to detect the passage of various sizes and types of flow disturbances (events) and to estimate their mean frequency of occurrence. Events characterized by large and sudden streamwise accelerations were found to be highly coherent throughout the log region and were strongly correlated with large fluctuations in wall shear-stress. Phase randomness between the near-wall quantities and the outer velocities was small. The results suggest that the flow events detected by conditional sampling applied to velocities in the log region may be related to the bursting process.

Robinson, S. K.↗

Utah FORGE: Well 16B(78)-32 Drilling Data

This drilling data for Utah FORGE well 16B(78)-32 include a well survey, core summary, mud and mud temperature logs, daily reports of the drilling process, and additional data from the Pason oil and gas company. Well 16B(78)-32 serves as the production well for reservoir creation, fluid circulation, and demonstration of heat extraction for the FORGE project. It has been drilled as a doublet approximately 300 feet parallel to and above the injection well 16A(78)-32. The proposed total depth was 10,658 feet, which was exceeded. Spudding began on April 26th, 2023. As of June 20th, 2023, the total depth measured 10,947 feet and the vertical depth measured 8,357 feet. Drilling included the trial use of insulated drill pipe (IDP) from Eavor Technologies, which was considered a success. IDP restricts counter-current heat transfer between drilling fluid inside the drill string and hotter returning fluid in the annulus, thereby ensuring the bottomhole assembly (BHA) remains submerged in cool fluid. Eavor rented 350 joints of IDP to FORGE which were run in two consecutive BHAs at the well. The results of this trial are included here.

15 GEOTHERMAL ENERGY↗

Towards a Comparative Assessment of Data-Driven Process Models in Health Information Technology

Process mining for conformance analysis focuses on comparing a reference process model against a data-driven process model that is generated via log files from information technology systems. While this approach is helpful when there is an existing process model in an organization, it leaves the question of what to do in the absence of a complete reference process model unanswered. In this paper, we present a comparative assessment approach that combines process mining, process mapping for dimensionality reduction, and statistical analysis. Our goal is to find similarities and dissimilarities in data-driven process models among U.S. Veterans Health Administration (VHA) facilities to assess process conformance among different healthcare facilities, which can help assess the standardization of care. We illustrate our approach by applying it to two clinical radiology order process models generated by two similar facilities. Our results demonstrate statistical similarities in the standardization of care among those two facilities.

Klasky, Hilda↗

A comparison of techniques for extracting emissivity information from thermal infrared data for geologic studies

This article evaluates three techniques developed to extract emissivity information from multispectral thermal infrared data. The techniques are the assumed Channel 6 emittance model, thermal log residuals, and alpha residuals. These techniques were applied to calibrated, atmospherically corrected thermal infrared multispectral scanner (TIMS) data acquired over Cuprite, Nevada in September 1990. Results indicate that the two new techniques (thermal log residuals and alpha residuals) provide two distinct advantages over the assumed Channel 6 emittance model. First, they permit emissivity information to be derived from all six TIMS channels. The assumed Channel 6 emittance model only permits emissivity values to be derived from five of the six TIMS channels. Second, both techniques are less susceptible to noise than the assumed Channel 6 emittance model. The disadvantage of both techniques is that laboratory data must be converted to thermal log residuals or alpha residuals to facilitate comparison with similarly processed image data. An additional advantage of the alpha residual technique is that the processed data are scene-independent unlike those obtained with the other techniques.

Hook, Simon J.↗

NREL Fleet Analysis Support Through Technology Integration Collaboration

This study leveraged the partnership between the United States Department of Energy's (DOE) Clean Cities Coalition Network and the Association for the Work Truck Industry (NTEA) to launch a vehicle and fleet analysis project that assisted fleets in identifying opportunities to save energy, improve efficiency, reduce costs, and meet environmental goals via short term data logging and analysis. The National Renewable Energy Laboratory (NREL) sought to establish a process that included initial data acquisition, provided data storage, and developed analytic methods to inform fleets of areas of opportunity based on approximately 30 days of in use vehicle performance data. However, long-term the project will require ongoing funding to fully develop and maintain the data sharing platform and to produce more complex analysis.

33 ADVANCED PROPULSION SYSTEMS↗

CVEVOLVE

CVEvolve is an agentic AI system for autonomous algorithm discovery for scientific data processing. It creates workflows where large language model agents freely set up and configure development environments and evaluation harnesses, develop and improve data processing algorithms with designed exploration-exploitation balancing mechanisms, log history and findings in a structured database, and run holdout testing to ensure algorithm generalizability. CVEvolve offers a zero-code interface and does not require users to provide structured data and evaluation scripts.

Cherukara, MatthewJoseph [Argonne National Laborat↗

Collecting and Processing Earth Science Data Metrics at NASA ESDIS

Since the launch of Terra satellite in 1999, the number of Earth Science remote sensing data products created and distributed by NASA's Earth Observing System (EOS) Data and Information System (EOSDIS) has increased from a few hundred to nearly ten thousand. NASA's Earth Science Data and Information System (ESDIS) Metrics System (EMS) collects metrics on data ingest, archive, and distribution by its Distributed Active Archive Centers (DAACs) and the Science Investigator-led Systems (SIPS), known as Data Providers. These metrics are critical in helping NASA management as well as data producers in resource planning and gaining a wide range of knowledge of data users and data usage.EMS receives flat files, or log files of data archive, ingest, and distribution either in their raw format, such as Apache web logs, or text files of log records formatted by the Data Providers. Tens of millions of records are processed each day to extract metrics on data products, user information, distribution protocols and services, and so on. The metrics are then made available to designated parties.This presentation provides an overview of the EMS processing workflow and improvement efforts made in recent years to handle ever-increasing number of data records and new metrics requirements, discusses several key steps including mapping log records to data products and identifying user communities along with geo-distribution, and demonstrates typical metrics capabilities produced by the EMS system. Challenges and potential approaches to improve the system are also discussed.

Pan, Jianfu↗

AD – Elog Data Search Using Natural Language Processing Techniques

The goal of this project is to develop and evaluate a Machine Learning model using natural language processing techniques to categorize AD E-Log entries efficiently. AD E-Logs, which play a vital role in Fermilab's projects as a system to record and manage various data, often suffer from slow search processing and imprecise results. To address this issue, we preprocessed the data and employed the Doc2Vec model for training. The resulting model enables us to identify and retrieve the most relevant entries, thus enhancing user experience in accessing desired information.

Muse, Amiin↗

Utah FORGE: 16B(78)-32 RFS DSS Strain Change Rate vs. Depth During 16A(78)-32 Stimulation

This dataset contains strain change rate versus depth data acquired using a Rayleigh frequency shift (RFS) distributed strain sensing (DSS) system during hydraulic stimulation of well 16A(78)-32 at the Utah FORGE site in April 2024. The data were collected from an optical fiber installed in the annulus of production well 16B(78)-32, approximately 300 feet from the injection well. The dataset includes tabulated strain data and an explanation of the methodology used to generate the frac log, which integrates strain change rate signals over selected time windows to identify fracture events.

15 GEOTHERMAL ENERGY↗

Structured Covariance Gaussian Networks for Orion Crew Module Aerodynamic Uncertainty Quantification

In this paper we propose a new approach for nonlinear regression and uncertainty quantification. The method is based on a pair of neural networks which parameterize mean and dense covariance functions of a multivariate Gaussian process, trained together to maximize the log-likelihood of observing the given data. The covariance matrix is made positive definite at every input by construction. We also propose a sampling approach that produces viable surrogate function realizations from the Gaussian process. We call the proposed model a Structured Covariance Gaussian Network (SCGN). We illustrate the use of SCGNs for learning an aerodynamic response surface with built-in uncertainty for the Orion crew module. We find that SCGN provides an efficient and systematic way to learn nonlinear functional relationships and dense covariances. We compare results to a baseline Gaussian process regressor and observe that the SCGN provides comparable uncertainty descriptions with improved scalability to dataset size. The sample functions generated by SCGN are fast to evaluate online and are therefore convenient for use in trajectory simulations. These results suggest that SCGN may be a viable computational method for aerodynamic uncertainty quantification.

machine learning↗

Structured Covariance Gaussian Networks for Orion Crew Module Aerodynamic Uncertainty Quantification

In this paper we propose a new approach for nonlinear regression and uncertainty quantification. The method is based on a pair of neural networks which parameterize mean and dense covariance functions of a multivariate Gaussian process, trained together to maximize the log-likelihood of observing the given data. The covariance matrix is made positive definite at every input by construction. We also propose a sampling approach that produces viable surrogate function realizations from the Gaussian process. We call the proposed model a Structured Covariance Gaussian Network (SCGN). We illustrate the use of SCGNs for learning an aerodynamic response surface with built-in uncertainty for the Orion crew module. We find that SCGN provides an efficient and systematic way to learn nonlinear functional relationships and dense covariances. We compare results to a baseline Gaussian process regressor and observe that the SCGN provides comparable uncertainty descriptions with improved scalability to dataset size. The sample functions generated by SCGN are fast to evaluate online and are therefore convenient for use in trajectory simulations. These results suggest that SCGN may be a viable computational method for aerodynamic uncertainty quantification.

machine learning↗

Skylab Medical Data Center and Archives

The founding of the Skylab medical data center and archives as a central area to house medical data from space flights is described. Skylab program strip charts, various daily reports and summaries, experiment reports and logs, status report on Skylab data quality, raw data digital tapes, processed data microfilm, and other Skylab documents are housed in the data center. In addition, this memorandum describes how the data center acted as a central point for the coordination of preflight and postflight baseline data and how it served as coordinator for all data processing through computation and analysis. Also described is a catalog identifying Skylab medical experiments and all related data currently archived in the data center.

Spross, F. R.↗