Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data reporting formats”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Classification and time of formation of Martian channels based on Viking data

The reported evaluation of Martian channel characteristics is based on Viking photographs taken from July 1976 to February 1977. The wide variation in crater densities shown by the considered Martian channels strongly implies widely differing ages for both fluviatile and lava channels. Attention is given to age determination methodology, a description of channels and implications for channel formation, surface water under present Martian conditions, surface water under more favorable Martian conditions in the past, channel parameter estimates, and volcanic channels.

Masursky, H.↗

MODIS Technical Report Series. Volume 4: MODIS data access user's guide: Scan cube format

The software described in this document provides I/O functions to be used with Moderate Resolution Spectroradiometer (MODIS) level 1 and 2 data, and could be easily extended to other data sources. This data is in a scan cube data format: a 3-dimensional ragged array containing multiple bands which have resolutions ranging from 250 to 1000 meters. The complexity of the data structure is handled internally by the library. The I/O calls allow the user to access any pixel in any band through 'C' structure syntax. The high MODIS data volume (approaching half a terabyte per day) has been a driving factor in the library design. To avoid recopying data for user access, all I/O is performed through dynamic 'C' pointer manipulation. This manual contains background material on MODIS, several coding examples of library usage, in-depth discussions of each function, reference 'man' type pages, and several appendices with details of the included files used to customize a user's data product for use with the library.

Kalb, Virginia L.↗

The virtual library: Coming of age

With the high speed networking capabilities, multiple media options, and massive amounts of information that exist in electronic format today, the concept of a 'virtual' library or 'library without walls' is becoming viable. In virtual library environment, the information processed goes beyond the traditional definition of documents to include the results of scientific and technical research and development (reports, software, data) recorded in any format or media: electronic, audio, video, or scanned images. Network access to information must include tools to help locate information sources and navigate the networks to connect to the sources, as well as methods to extract the relevant information. Graphical User Interfaces (GUI's) that are intuitive and navigational tools such as Intelligent Gateway Processors (IGP) will provide users with seamless and transparent use of high speed networks to access, organize, and manage information. Traditional libraries will become points of electronic access to information on multiple medias. The emphasis will be towards unique collections of information at each library rather than entire collections at every library. It is no longer a question of whether there is enough information available; it is more a question of how to manage the vast volumes of information. The future equation will involve being able to organize knowledge, manage information, and provide access at the point of origin.

Hunter, Judy F.↗

Soil microbial ecology and microbiome-metabolite linkages improve understanding of ecosystem states along terrestrial-aquatic interfaces

These data are from Bandopadhyay et al., "Soil microbial ecology and microbiome-metabolite linkages improve understanding of ecosystem states along terrestrial-aquatic interfaces". This study aims to understand the soil microbial ecology along terrestrial-aquatic interfaces of a freshwater and estuarine region and how it relates to organic matter. We analyzed soil microbial (16S rRNA gene) and organic matter (Fourier-transform ion cyclotron resonance mass spectrometry, FTICR-MS) composition from upland (forested), transition (stressed forest), and wetland positions at three sites in each of the Lake Erie (freshwater) and Chesapeake Bay (estuarine) regions. This dataset includes 16S rRNA gene amplicon data (only processed file types included here) and organic matter composition from FTICR-MS data (raw and processed files included here) from upland (forested), transition (stressed forest), and wetland positions at three sites in each of the Lake Erie and Chesapeake Bay regions. These sites are part of the COMPASS-FME project (https://compass.pnnl.gov/FME/COMPASSFME). File formats and software needed to access files: 16S rRNA gene amplicon data: These files follow the format reported here https://ess-dive.gitbook.io/amplicon-sequencing-reporting-format#updates-in-v1.0.1. As per this format, there are four file types reported: 1. Taxon tables (also called sequence-by-sample or OTU (operational taxonomic unit)/ESV (exact sequence variant) tables) : available in a .txt file format and accessible using TextEdit or MS Excel. 2. Representative sequences (also called consensus sequences) : available in a .fasta format and accessible using TextEdit. 3. Sequencing metadata : available in a MS Excel workbook file format and CSV file format 4. Bioinformatic metadata : available in a MS Excel workbook file format and CSV file format FTICR-MS data: 1. Raw data converted to a processed file with intensities of the peaks in the given samples : available in a MS Excel CSV file format 2. Processed file used in analyses and visualizations (appended as icr_long_) : available in a MS Excel CSV file format 3. Metadata file for ICR features (appended as icr_meta) : available in a MS Excel CSV file format

54 ENVIRONMENTAL SCIENCES↗

Leaf gas exchange and fitted parameters, two sites in Panama, 2022

Photosynthetic CO2 response curves (ACi curves), light response curves (AQ curves), dark adapted dark respiration, conductance curves, and survey measurements for leaves measured in the Parque Natural Metropolitano (PNM) and Gamboa, Panama, from January to April 2022 are presented. Measurements were made on leaves from 46 different tree, shrub and liana species, from sunlit canopy and understory locations. Leaf area index from 8 vertical profiles at PNM are also included. The aim of this measurement campaign was three-fold: (i) to improve our understanding of the vertical variation in leaf-level water use efficiency; (ii) to develop an understanding of the sensitivity of stomata to changes in environmental conditions, especially light, humidity, and temperature; (iii) to improve models which can predict leaf traits from leaf contact spectral measurements. All the gas exchange data and metadata are presented in .csv files and complete instrument output are included in .zip folders. Data and metadata meet the ESS-DIVE leaf-level gas exchange reporting format requirements. The protocol details are provided as pdf documents. In addition to gas exchange data reported here these samples were also used for measurement of leaf optical properties, carbon and nitrogen content, and leaf mass per unit leaf area (LMA). These data can be cross linked using the unique sample ID and are provided in separate data packages.

54 ENVIRONMENTAL SCIENCES↗

Groundwater and river water elevations and temperature from 2017 to 2022 across Meander Z in the East River Watershed, Colorado

This dataset includes groundwater and river water elevations and temperature data collected in the East River watershed located in the Upper Colorado River Basin. The data were collected in order to investigate the coupling between hydrology and biogeochemical processes in the floodplain. Data was collected at ten groundwater locations in Meander Z (MZ), located just upstream of the confluence with Brush Creek and two river locations directly adjacent to Meander Z from 2017-2019. From 2019-2022, data was collected at five groundwater locations in Meander Z. Note that location names, not location identifiers (IDs), are used in the related publication Dewey et al. (2022). Both location IDs and names are included in data files. Files in this dataset include the main data files for each location zipped into a single folder (waterlevel_data.zip), an installation methods file describing sensor installation (InstallationMethods.csv), a file containing field metadata including GPS (Global Positioning System) coordinates and ground surface elevations (transducers_locations.csv). This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type. This dataset conforms to the ESS-DIVE hydrological reporting format. 2026-04-27 Update: The river water elevation data files (ER-MZR1.csv and ER-MZR2.csv) were corrected. The data for these two locations were inadvertently swapped in the original published data. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

54 ENVIRONMENTAL SCIENCES↗

ESS-DIVE Reporting Format for File-level Metadata

The ESS-DIVE reporting format for file-level metadata (FLMD) provides granular information at the data file level to describe the contents, scope, and structure of the data file to enable comparison of data files within a data package. The FLMD are fully consistent with and augment the metadata collected at the data package level. We developed the FLMD template based on a review of a small number of existing FLMD in use at other agencies and repositories with valuable input from the Environmental Systems Science (ESS) Community. Also included is a template for a CSV Data Dictionary where users can provide file-level information about the contents of a CSV data file (e.g., define column names, provide units). Files are in .csv, .xlsx, and .md. Templates are in both .csv and .xlsx (open with e.g. Microsoft Excel, LibreOffice, or Google Sheets). Open the .md files by downloading and using a text editor (e.g. Notepad or TextEdit). Though we provide Excel templates for the file-level metadata reporting format, our instructions encourage users to 'Save the FLMD template as a CSV following the CSV Reporting Format guidance'. In addition, we developed the ESS-DIVE File Level Metadata Extractor which is a lightweight python script that can extract some FLMD fields following the recommended FLMD format and structure.

54 ENVIRONMENTAL SCIENCES↗

Software to Compare NPP HDF5 Data Files

This software was developed for the NPOESS (National Polar-orbiting Operational Environmental Satellite System) Preparatory Project (NPP) Science Data Segment. The purpose of this software is to compare HDF5 (Hierarchical Data Format) files specific to NPP and report whether the HDF5 files are identical. If the HDF5 files are different, users have the option of printing out the list of differences in the HDF5 data files. The user provides paths to two directories containing a list of HDF5 files to compare. The tool would select matching HDF5 file names from the two directories and run the comparison on each file. The user can also select from three levels of detail. Level 0 is the basic level, which simply states whether the files match or not. Level 1 is the intermediate level, which lists the differences between the files. Level 2 lists all the details regarding the comparison, such as which objects were compared, and how and where they are different. The HDF5 tool is written specifically for the NPP project. As such, it ignores certain attributes (such as creation_date, creation_ time, etc.) in the HDF5 files. This is because even though two HDF5 files could represent exactly the same granule, if they are created at different times, the creation date and time would be different. This tool is smart enough to ignore differences that are not relevant to NPP users.

Wiegand, Chiu P.↗

Computed Tomography Scanning and Geophysical Measurements of Appalachian Basin Core from the Jones and Laughlin #1 Well, Beaver County, PA

The computed tomography (CT) facilities and the Multi-Sensor Core Logger (MSCL) at the National Energy Technology Laboratory (NETL) in Morgantown, West Virginia, were used to characterize Appalachian Basin core from Beaver County, Pennsylvania. The primary impetus of this work is a collaboration between the U.S. Department of Energy (DOE) and the Pennsylvania Geological Survey to characterize and make publicly available core information from the Onondaga-Huntersville formations of the Appalachian Basin. This stratigraphic well and the core data produced in this report will aid in understanding the structural complexities of the Onondaga-Huntersville formations. The resultant datasets are presented in this report and can be accessed from NETL's Energy Data eXchange (EDX) online system using the following link: https://edx.netl.doe.gov/dataset/jonesandlaughlin1well. All equipment and techniques used were non-destructive, enabling future examinations and analyses to be performed on these cores. Fractures, discontinuities, and millimeter-scale features were readily detectable with imaging performed with the NETL medical CT scanner over the entire core. Qualitative analysis of the medical CT images, coupled with X-ray fluorescence (XRF), and magnetic susceptibility measurements from the MSCL were useful in identifying zones of interest for further study. Targeted higher resolution CT scanning of select sections was performed with NETL’s micro-CT scanner. The combination of methods used provides a multiscale analysis of the core; the resulting macro and micro descriptions are relevant to many subsurface energy related examinations traditionally performed at NETL.

58 GEOSCIENCES↗

Leveraging large language models to address data scarcity in machine learning for graphene synthesis

Machine learning in experimental materials science faces significant challenges due to the scarcity of data, which are costly and time-consuming to generate, particularly when relying on in-house experiments. Literature data mining offers a potential solution but introduces issues like mixed data quality, inconsistent formats, and non-uniform reporting of synthesis parameters, resulting in partially missing and heterogeneous features across the dataset. Here, we propose data imputation and feature engineering methods that employ pre-trained large language models (LLMs) to enhance machine learning performance on scarce, heterogeneous datasets, demonstrated on graphene CVD synthesis data and the ML-HydPARK hydrogen storage dataset. GPT models perform data imputation via tailored prompting and semantic normalization of inconsistently reported features through embeddings, for example, to harmonize the complex nomenclature of CVD substrates. Beyond yielding more diverse and richer feature representations than traditional methods such as K-nearest neighbors (KNN) and Multivariate Imputation by Chained Equations (MICE), LLM-based data imputation is evaluated against dataset characteristics and prompting strategies. We vary the level of autonomy granted to the LLM, from generic prompting that leverages pre-trained knowledge for autonomous data generation to data-informed prompting that constrains outputs using target-specific information, and demonstrate which level of autonomy yields superior imputation performance across datasets and feature types. The proposed data engineering methods markedly improve downstream performance; for example, in graphene layer number classification using a support vector machine (SVM), binary accuracy increases from 39% to 65% and ternary accuracy from 52% to 72%. Fine-tuning experiments on both datasets show that combining our proposed LLM-based data imputation and feature encoding methods with numerical machine learning predictors outperforms standalone fine-tuned LLM predictors in data-scarce settings. The proposed strategies emphasize data enhancement techniques rather than refining learning architectures or regularizing loss functions, offering a broadly applicable framework for improving machine learning performance on scarce, inhomogeneous datasets.

Chemical vapor deposition↗

Integrated Hourly Meteorological Database of 20 Meteorological Stations (1981-2022) for Watershed Function SFA Hydrological Modeling

This dataset contains (a) a script “R_met_integrated_for_modeling.R”, and (b) associated input CSV files: 3 CSV files per location to create a 5-variable integrated meteorological dataset file (air temperature, precipitation, wind speed, relative humidity, and solar radiation) for 19 meteorological stations and 1 location within Trail Creek from the modeling team within the East River Community Observatory as part of the Watershed Function Scientific Focus Area (SFA). As meteorological forcings varied across the watershed, a high-frequency database is needed to ensure consistency in the data analysis and modeling. We evaluated several data sources, including gridded meteorological products and field data from meteorological stations. We determined that our modeling efforts required multiple data sources to meet all their needs. As output, this dataset contains (c) a single CSV data file (*_1981-2022.csv) for each location (20 CSV output files total) containing hourly time series data for 1981 to 2022 and (d) five PNG files of time series and density plots for each variable per location (100 PNG files). Detailed location metadata is contained within the Integrated_Met_Database_Locations.csv file for each point location included within this dataset, obtained from Varadharajan et al., 2023 doi:10.15485/1660962. This dataset also includes (e) a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and (f) a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type. Review the (g) ReadMe_Integrated_Met_Database.pdf file for additional details on the script, methods, and structure of the dataset.The script integrates Northwest Alliance for Computational Science and Engineering’s PRISM gridded data product, National Oceanic and Atmospheric Administration’s NCEP-NCAR Reanalysis 1 gridded data product (through the `RCNEP` R package, Kemp et al., doi:10.32614/CRAN.package.RNCEP), and analytical-based calculations. Further, this script downscales the input data into hourly frequency, which is necessary for the modeling efforts.

54 ENVIRONMENTAL SCIENCES↗

MINIS: Multipurpose Interactive NASA Information System

The Multipurpose Interactive NASA Information Systems (MINIS) was developed in response to the need for a data management system capable of operation on several different minicomputer systems. The desired system had to be capable of performing the functions of a LANDSAT photo descriptive data retrieval system while remaining general in terms of other acceptable user definable data bases. The system also had to be capable of performing data base updates and providing user-formatted output reports. The resultant MINI System provides all of these capabilities and several other features to complement the data management system. The MINI System is currently implemented on two minicomputer systems and is in the process of being installed on another minicomputer system. The MINIS is operational on four different data bases.

Source record↗

Preliminary Assessment Of The Burning Dynamics Of Jp8 Droplets In Microgravity

In this report we present new data for fuel droplet combustion in microgravity to examine the influence of ambient gas and fuel composition on flame structure and sooting dynamics for droplets with initial diameters in the range of 0.4mm to 0.5mm. The fuels are JP8 (a kerosene derivative) and nonane. The ambient gas is air and a mixture of 30% oxygen and 70% helium, the latter having been examined for burning under conditions where soot formation is minimal. Some data at elevated pressures are also reported. The burning process shows a nonlinear D2 progression which is independent of soot formation as burning in a helium inert showed the same nonlinear trend. Flames were proportionally farther from the droplet surface in helium than they were in air. A nondimensional parameter is presented that consolidates the three standoff distances for the droplet, flame and soot shell diameters within the initial diameter ranges examined.

Bae, J. H.↗

Aerobic respiration controls on shale weathering, Geochimica et Cosmochimica Acta, 2023: Dataset

This data package was generated in order to support the development of a deep-time weathering model and to assess the coupling between shale weathering and aerobic respiration in the paper “Aerobic respiration controls on shale weathering” by Stolze et al., Geochimica et Cosmochimica Acta (2023). The package contains two csv files providing the average CO2(g) concentration profiles [ppm] and mineral concentration profiles [wt%], respectively. The CO2(g) concentration profiles were measured in the vicinity of the monitoring well PLM2 between January 2018 and April 2019. The gas samples were collected in the unsaturated zone to a depth of 1.52 m. The mineral concentration profiles were determined by X-Ray diffraction (XRD). The XRD measurements were performed on sub-core samples collected in the monitoring well PLM3 down to a depth of 7.01 m. The dataset additionally includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata; and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.Update on 2024-05-28: Revised versions of the CSV data files (CO2_data_GCA_Stolze_et_al_2023.csv and XRD_data_GCA_Stolze_et_al_2023.csv) were made to apply ESS-DIVE's CSV reporting format guidelines. Updated versions of the File Level Metadata (v2_20240528_flmd.csv) and Data Dictionary (v2_20240528_dd.csv) files were updated to reflect the changes made to the CSV files.

54 ENVIRONMENTAL SCIENCES↗

SG50 Data-format Requirement Document for an Automatically Readable, Comprehensive and Curated Experimental Reaction Database MEDUSA

This report constitutes the requirement document that guides the development of the experimental reaction database, MEDUSAL (Machine-readable Experimental Data User App & Library), created by OECD/NEA/WPEC SubGroup 50. Experimental reaction data are usually stored in the EXFOR library in EXFOR format. With MEDUSAL, the WPEC sub-group 50 wants to go beyond the EXFOR format and database to generate a library that is (a) automatically readable, (b) comprehensive, and (c) curated.

Nuclear Criticality Safety Program (NCSP)↗

COMPASS-FME Terrestrial Ecosystem Manipulation to Probe the Effects of Storm Treatments (TEMPEST) Experiment Level 1 Sensor Data v1-2

This is the version 1-2 Level 1 (L1) data release for COMPASS-FME environmental sensors located at our Terrestrial Ecosystem Manipulation to Probe the Effects of Storm Treatments (TEMPEST) experimental site. This manipulative, ecosystem-scale TEMPEST experiment addresses the potential for freshwater and estuarine-water disturbance events to alter tree function, species composition, and ecosystem processes in a deciduous coastal forest in MD, USA. The experiment uses a large-unit (2000 m2), un-replicated experimental design, with three 50 m × 40 m plots serving as control, freshwater, and estuarine-water treatments.L1 data are close to raw, but are units-transformed and have out-of-instrument-bounds and out-of-service flags added. Duplicates and missing data are removed but otherwise these data are not filtered, and have not been subject to any additional algorithmic or human QA/QC. Any scientific analyses of L1 data should be performed with care. **This dataset will be updated quarterly with new data for the duration of the project**This dataset includes:- An overall dataset README file that describes the current version, gives citation and contact information, etc.- Site- and year-specific folders, each holding up to 12 CSV (comma separated value) data files for each site and plot in that year.- Metadata files within each site-year folder provide full information on data units, expected ranges, contact information, detailed flood times, as well as a general description of the site.- Environmental sensor types that appear in the data files include weather (ClimaVUE50, CS, RM Young, and LI instruments in the graphs below); soil conditions (TEROS12); soil redox state (Redox); groundwater variables (AquaTROLL200 and AquaTROLL600); open water sondes (Exo); tree sap velocity (Sapflow); and system voltage and state (Datalogger). Data are normally logged every 15 minutes.Please see v1-2 TEMPEST L1 Sensor Package Quick Start.pdf for detailed information on data package structure, temporal coverage, and versioning.The TEMPEST flood events occurred on the following dates. They lasted for ~10 hours each day and delivered ~80,000 gallons to each plot; many data streams are available at 1 or 5 minute frequency during these periods.* Tests: Aug 25 (fresh plot) and Sep 9 (salt plot), 2021* TEMPEST 1: June 22, 2022* TEMPEST 2: June 6-7, 2023* TEMPEST 3: June 11-13, 2024

54 ENVIRONMENTAL SCIENCES↗