Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “open data format”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

CHESS 2025: Waveform LiDAR data from NEON AOP surveys

This dataset provides Level 1 (L1) full-waveform light detection and ranging (LiDAR) data collected for the 2025 Colorado Headwaters Ecological Spectroscopy Study (CHESS). These data were acquired to enable characterization of vegetation structure and other three-dimensional features of the land surface, and to evaluate structural changes that may have occurred between a prior LiDAR acquisition in 2018 and the 2025 overflight. Waveform LiDAR data can provide more detailed information about objects on the ground than discrete point clouds typically do, and they are often used for granular target segmentation and characterization of subcanopy vegetation. The data were acquired over three study domains in the Upper Gunnison river basin: the upper East River watershed (CRBU); Almont Triangle and Taylor Canyon (ALMO); and Upper Taylor River watershed (UPTA) between 2025-06-13 and 2025-07-15. LiDAR data were acquired using the Optech Galaxy Prime Airborne LiDAR Terrain Mapper onboard the National Ecological Observatory Network (NEON) Airborne Observation Platform (AOP). These are the primary waveform LiDAR data delivered by NEON and are provided per flightline in compressed Pulsewaves format, an open-source binary file standard. A Pulsewaves object comprises a two files: a pulse (.pls) file, which stores the geographic origin, outgoing vector, and metadata for every laser pulse emitted by the scanner, and a wave file (.wvs), which stores the sequential amplitude samples of the outgoing pulse and the returning signals. The files are published here in their compressed forms (.plz, .wvz). All waveform data were processed following the theoretical workflow described in the NEON L0-to-L1 Waveform LiDAR Algorithm Theoretical Basis Document (Krause and Goulden 2022a); however, the Pulsewaves output format differs from a legacy format described in that document. Waveform amplitude samples are recorded at 1 nanosecond intervals. All coordinates are provided in meters. Horizontal coordinates are referenced in Universal Transverse Mercator (UTM) zone 13N and the World Geodetic System (WGS) 1984 ensemble datum. Elevations are referenced to Geoid12A. Waveform data for the UPTA survey area were collected without incident and the published records are complete. However, both the ALMO and CRBU collections experienced issues that resulted in incomplete data for those areas. On collection day 2018-06-16 a hardware failure caused the waveform digitizer to lose data from the eastern edge of the ALMO site (Figure 22). The waveform data for flightlines 2–20 could not be extracted from the digitizer, and the data proved unrecoverable. As a result, a portion of the site does not have coverage with waveform data. Although no hardware failure was observed during collection over the CRBU area, final waveform files generated by vendor software contained only ~25% of the expected number of return pulses. After discovery, NEON initiated troubleshooting with the vendor. The root cause of the data ablation had not been identified at the time of publication. Additional data will be published in an update to this package if further recovery proves successful. CHESS Project Description: The Colorado Headwaters Ecological Spectroscopy Study (CHESS) comprised a multi-week airborne remote sensing and field observation campaign in the Upper Gunnison Basin, Colorado, conducted in June and July of 2025. Airborne remote sensing was conducted by the National Ecological Observatory Network Airborne Observation Platform (NEON AOP), concurrent with a field campaign run by the Rocky Mountain Biological Laboratory (RMBL), the Lawrence Berkeley National Laboratory (LBNL) and SLAC National Accelerator Laboratory Watershed Function Science Focus Area (SFA), and NASA-JPL (Jet Propulsion Laboratory) Earth Surface Mineral Dust Source Investigation (EMIT) program. Between June 10 and July 18, 2025, the NEON AOP flight team collected high-resolution aerial imaging spectroscopy and Light Detection and Ranging (LiDAR) data over three domains: the Upper East River (CRBU), Almont Triangle (ALMO), and the Upper Taylor Basin (UPTA). In coordination with the flights, a field campaign acquired ground-truth observations, including observations of vegetation composition, foliar traits, forest demography, and subsurface properties in 18 core sampling areas within the domains. Additional surface water observations were taken at over 380 point locations. All CHESS campaign datasets can be found within the CHESS ESS-DIVE data portal: https://data.ess-dive.lbl.gov/portals/chess. Funding Acknowledgement: Field and remote-sensing data acquisition was performed under a grant from the National Aeronautics and Space Administration (80NSSC24K1005). This work was also supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

2018 NEON and 2025 CHESS Campaigns↗

GES DISC Greenhouse Gas Data Sets and Associated Services

NASA Goddard Earth Sciences (GES) Data and Information Services Center (DISC) archives and distributes rich collections of data on atmospheric greenhouse gases from multiple missions. Hosted data include those from the Atmospheric Infrared Sounder (AIRS) mission (which has observed CO2, CH4, ozone, and water vapor since 2002); legacy water vapor and ozone retrievals from TIROS Operational Vertical Sounder (TOVS); and Upper Atmosphere Research Satellite (UARS) going back to the early 1980s. GES DISC also archives and supports data from seven projects of the Making Earth System Data Records for Use in Research Environments (MEaSUREs) program that have ozone and water vapor records. Greenhouse gases data from the A-Train satellite constellation is also available: (1) Aura-Ozone Monitoring Instrument (OMI) and Microwave Limb Sounder (MLS) ozone, nitrous oxide, and water vapor since 2004; (2) Greenhouse Gases Observing Satellite (GOSAT) CO2 observations since 2009 from the Atmospheric CO2 Observations from Space (ACOS) task; and (3) Orbiting Carbon Observatory-2 (OCO-2) CO2 data since 2014. The most recent related data set that the GES DISC archives is methane flux for North America, as part of NASAs Carbon Monitoring System (CMS) project. This dataset contains estimates of methane emission in North America based on an inversion of the GEOS-Chem chemical transport model constrained by GOSAT observations (Turner et al., 2015). Along with data stewardship, an important focus area of the GES DISC is to enhance the usability of its data and broaden its user base. Users have unrestricted access to a new user-friendly search interface, which includes many services such as variable subsetting, format conversion, quality screening, and quick browse. The majority of the GES DISC data sets are also accessible through Open-source Project for a Network Data Access Protocol (OPeNDAP) and Web Coverage Service (WCS). The latter two services provide more options for specialized subsetting, format conversion, and image viewing. Additional data exploration, data preview, and preliminary analysis capabilities are available via NASA Giovanni, which obviates the need forusers to download the data (Acker and Leptoukh, 2007). Giovanni provides a bridge between the data and science and has been very successful in extending GES DISC data to educational users and to users with limited resources.

data ordering↗

Automating Traffic Microsimulation from SYNCHRO UTDF to SUMO

Modern transportation research relies on seamlessly integrating traffic signal data with robust network representation and simulation tools. This study presents utdf2gmns, an open-source Python tool that automates conversion of the Universal Traffic Data Format, including network representation, signalized intersections, and turning volumes into the General Modeling Network Specification (GMNS) Standard. The resulting GMNS-compliant network can be converted for microsimulation in SUMO. By automatically extracting intersection control parameters and aligning them with GMNS conventions, utdf2gmns minimizes manual preprocessing and data loss. utdf2gmns also integrates with the Sigma-X engine to extract and visualize key traffic control metrics, such as phasing diagrams, turning volumes, volume-tocapacity ratios, and control delays. This streamlined workflow enables efficient scenario testing, accurate model building, and consistent data management. Validated through case studies, utdf2gmns reliably models complex urban corridors, promoting reproducibility and standardization. Documentation is available on GitHub and PyPI, supporting easy integration and community engagement.

Luo, Roy [ORNL] (ORCID:0009000312909983)↗

Provenance in Data Interoperability for Multi-Sensor Intercomparison

As our inventory of Earth science data sets grows, the ability to compare, merge and fuse multiple datasets grows in importance. This requires a deeper data interoperability than we have now. Efforts such as Open Geospatial Consortium and OPeNDAP (Open-source Project for a Network Data Access Protocol) have broken down format barriers to interoperability; the next challenge is the semantic aspects of the data. Consider the issues when satellite data are merged, cross-calibrated, validated, inter-compared and fused. We must match up data sets that are related, yet different in significant ways: the phenomenon being measured, measurement technique, location in space-time or quality of the measurements. If subtle distinctions between similar measurements are not clear to the user, results can be meaningless or lead to an incorrect interpretation of the data. Most of these distinctions trace to how the data came to be: sensors, processing and quality assessment. For example, monthly averages of satellite-based aerosol measurements often show significant discrepancies, which might be due to differences in spatio- temporal aggregation, sampling issues, sensor biases, algorithm differences or calibration issues. Provenance information must be captured in a semantic framework that allows data inter-use tools to incorporate it and aid in the intervention of comparison or merged products. Semantic web technology allows us to encode our knowledge of measurement characteristics, phenomena measured, space-time representation, and data quality attributes in a well-structured, machine-readable ontology and rulesets. An analysis tool can use this knowledge to show users the provenance-related distrintions between two variables, advising on options for further data processing and analysis. An additional problem for workflows distributed across heterogeneous systems is retrieval and transport of provenance. Provenance may be either embedded within the data payload, or transmitted from server to client in an out-of-band mechanism. The out of band mechanism is more flexible in the richness of provenance information that can be accomodated, but it relies on a persistent framework and can be difficult for legacy clients to use. We are prototyping the embedded model, incorporating provenance within metadata objects in the data payload. Thus, it always remains with the data. The downside is a limit to the size of provenance metadata that we can include, an issue that will eventually need resolution to encompass the richness of provenance information required for daata intercomparison and merging.

Lynnes, Chris↗

pyEGAF: An open-source Python library for the Evaluated Gamma-ray Activation File

The Evaluated Gamma-ray Activation File (EGAF) is one of the most comprehensive resources for thermal neutron-capture data. This database contains data from prompt gamma activation analysis measurements carried out in a consistent manner using the same experimental configuration at the Budapest Research Reactor for 245 isotopes. Although these valuable datasets have been freely available for many years, one of the drawbacks is the outdated and cryptic Evaluated Nuclear Structure Data File (ENSDF) format that is currently adopted for dissemination, making it difficult for users unfamiliar with the format to access and utilize the data contained therein. Furthermore, the ENSDF format does not readily lend itself to modern computational technologies and a parser is required to interpret the complicated mixed-record format. To help overcome these challenges, we have developed a translator to convert the ENSDF-formatted datasets into an open standard JavaScript Object Notation (JSON) format enabling accessibility to applications using different programming languages running in different environments. To compliment this effort, we have also developed an open-source software package implemented in Python, pyEGAF, that is designed to interact with the JSON data structures for general purpose access, manipulation, and analysis of the neutron-capture $\gamma$-ray data in EGAF. The new format, together with the pyEGAF library, greatly enhances access to the wider applications community where EGAF data may be useful or is required.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Common Data Format (CDF) and Coordinated Data Analysis Web (CDAWeb)

The Coordinated Data Analysis Web (CDAWeb) data browsing system provides plotting, listing and open access v ia FTP, HTTP, and web services (REST, SOAP, OPeNDAP) for data from mo st NASA Heliophysics missions and is heavily used by the community. C ombining data from many instruments and missions enables broad resear ch analysis and correlation and coordination with other experiments a nd missions. Crucial to its effectiveness is the use of a standard se lf-describing data format, in this case, the Common Data Format (CDF) , also developed at the Space Physics Data facility , and the use of metadata standa rds (easily edited with SKTeditor ). CDAweb is based on a set of IDL routines, CDAWlib . . The CDF project also maintains soft ware and services for translating between many standard formats (CDF. netCDF, HDF, FITS, XML) <! .

Candey, Robert M.↗

IECM calibration and data reduction requirements

The induced environment contamination monitor (IECM) tape recorder format, as it relates to the ouput of meaningful data from the IECM instrument, is explained. Eight-bit words (or bytes) generate numbers that represent voltage levels of electronic detection probes for each experiment. This information is amalgamated by the IECM Data Acquisition and Control System (DACS). In some cases bits represent certain status situations concerning an experiment, such as whether a valve is opened or closed. Voltages are transformed into meaningful physical phenomena through equations of calibration. Data formats and plots are generated as requested for each IECM experimenter.

Wills, F. D.↗

NASA GES DISC support of CO2 Data from OCO-2, ACOS, and AIRS

NASA Goddard Earth Sciences Data and Information Services Centers (GES DISC) is the data center assigned to archive and distribute current AIRS, ACOS data and data from the upcoming OCO-2 mission. The GES DISC archives and supports data containing information on CO2 as well as other atmospheric composition, atmospheric dynamics, modeling and precipitation. Along with the data stewardship, an important mission of GES DISC is to facilitate access to and enhance the usability of data as well as to broaden the user base. GES DISC strives to promote the awareness of science content and novelty of the data by working with Science Team members and releasing news articles as appropriate. Analysis of events that are of interest to the general public, and that help in understanding the goals of NASA Earth Observing missions, have been among most popular practices.Users have unrestricted access to a user-friendly search interface, Mirador, that allows temporal, spatial, keyword and event searches, as well as an ontology-driven drill down. Variable subsetting, format conversion, quality screening, and quick browse, are among the services available in Mirador. The majority of the GES DISC data are also accessible through OPeNDAP (Open-source Project for a Network Data Access Protocol) and WMS (Web Map Service). These services add more options for specialized subsetting, format conversion, image viewing and contributing to data interoperability.

Wei, Jennifer C↗

cjohnson-LANL/GRL_Kilauea

The python routines are outlined in detail to perform the methods and results in the manuscript under review in the journal Geophysical Research Letters titled “Seismic features predict ground motions during repeating caldera collapse sequence” with LA-UR-23-33345. All routines are written in open source python and were applied to publicly available data sets. The codes formats the data into the appropriate structure required to train a boosted tree regression model. Other codes produce figure results.

Johnson, Christopher W↗

AI for Earthquake Physics

The core LANL program sponsored by Office of Science, Basic Energy Science, Chemical Sciences, Geosciences, and Biosciences (DOE-BES-CSGB) and led by PI Johnson aims to research earthquake faults to advance fault physics and earthquake hazards. All work completed is required to be made publicly available through publications and open-source codes supporting the published results. All routines are/will-be written in open source python and applied to publicly available data sets. These routines will format data from input into models, develop and test modeling frameworks for the problems addressed, and produce figures applicable to peer-reviewed manuscripts. All work is reviewed for Los Alamos Unlimited Release before submitting to a journal. This summary encompasses recently completed work and work to be complete for the duration of the program.

Johnson, Christopher↗

Vaporization response of evaporating drops with finite thermal conductivity

A numerical computing procedure was developed for calculating vaporization histories of evaporating drops in a combustor in which travelling transverse oscillations occurred. The liquid drop was assumed to have a finite thermal conductivity. The system of equations was solved by using a finite difference method programmed for solution on a high speed digital computer. Oscillations in the ratio of vaporization of an array of repetitivity injected drops in the combustor were obtained from summation of individual drop histories. A nonlinear in-phase frequency response factor for the entire vaporization process to oscillations in pressure was evaluated. A nonlinear out-of-phase response factor, in-phase and out-of-phase harmonic response factors, and a Princeton type 'n' and 'tau' were determined. The resulting data was correlated and is presented in graphical format. Qualitative agreement with the open literature is obtained in the behavior of the in-phase response factor. Quantitatively the results of the present finite conductivity spray analysis do not correlate with the results of a single drop model.

Agosta, V. D.↗

Biolink Model: A universal schema for knowledge graphs in clinical, biomedical, and translational science

Abstract Within clinical, biomedical, and translational science, an increasing number of projects are adopting graphs for knowledge representation. Graph‐based data models elucidate the interconnectedness among core biomedical concepts, enable data structures to be easily updated, and support intuitive queries, visualizations, and inference algorithms. However, knowledge discovery across these “knowledge graphs” (KGs) has remained difficult. Data set heterogeneity and complexity; the proliferation of ad hoc data formats; poor compliance with guidelines on findability, accessibility, interoperability, and reusability; and, in particular, the lack of a universally accepted, open‐access model for standardization across biomedical KGs has left the task of reconciling data sources to downstream consumers. Biolink Model is an open‐source data model that can be used to formalize the relationships between data structures in translational science. It incorporates object‐oriented classification and graph‐oriented features. The core of the model is a set of hierarchical, interconnected classes (or categories) and relationships between them (or predicates) representing biomedical entities such as gene, disease, chemical, anatomic structure, and phenotype. The model provides class and edge attributes and associations that guide how entities should relate to one another. Here, we highlight the need for a standardized data model for KGs, describe Biolink Model, and compare it with other models. We demonstrate the utility of Biolink Model in various initiatives, including the Biomedical Data Translator Consortium and the Monarch Initiative, and show how it has supported easier integration and interoperability of biomedical KGs, bringing together knowledge from multiple sources and helping to realize the goals of translational science.

60 APPLIED LIFE SCIENCES↗

The MY NASA DATA Project

On the one hand, locating the right dataset, then figuring out how to use it, is a daunting task that is familiar to almost any scientist or graduate student in the fields of Earth system science. On the other hand, the ability to explore authentic Earth system science data, through inquiry-based education, is an important goal in US national education standards. Fortunately, in the digital age, tools are emerging that can make such data exploration commonplace at all educational levels. This paper describes the conception and development of one project that aims to bridge this gap: Mentoring and inquiry using NASA Data on Atmospheric and Earth science for Teachers and Amateurs (MY NASA DATA; mynasadata.larc.nasa.gov). With funding from NASA's Science Mission Directorate, this project was launched in early 2004 with the aim of developing microsets and identifying other enablers for making data accessible. A key feature of the project is a Live Access Server, the first educational implementation of this open source software, developed by NOAA, that makes it possible to explore multiple data formats through a single interface. This powerful tool is made more useful to the primary target audiences (K-12 and amateur scientists) through careful selection of the data offered, user-friendly explanations of the tool itself, and age-appropriate explanations of the parameters. However experience already shows that graduate students and even practicing scientists can also make use of this resource. The website also hosts teacher-contributed lesson plans, and seeks to feature reports of research projects that use the data.

Chambers, Lin H.↗

Hail Storm Risk Assessment Using Space-Borne Remote Sensing Observations and Reanalyses

Much of the world is impacted by severe thunderstorms, but whether they become disasters depends upon resilience--our capacity to prepare, mitigate, respond, and recover. Hail is the costliest severe weather hazard for the insurance industry, generating ~70% of severe convective storm losses due to damage to assets such as homes, businesses, agriculture, and infrastructure. Most insurance companies do not reserve enough capital to cover catastrophes, so they acquire reinsurance. The reinsurance industry uses catastrophe models (CatModels) to statistically estimate risk to an insurer’s portfolio. Hail CatModels are developed with climatologies that define hailstorm frequency and severity. Hail-prone areas can be defined using hail reports from trained spotters, the media, and the general public. Extremely severe hail (2+ inch diameter) occurs nearly every day across the world. Weather radars can detect hail because hailstones strongly reflect microwave signals that they emit. However, hail climatologies are difficult to derive because hail covers small areas and there are neither hail reporting mechanisms (e.g. website or mobile app) nor radar networks in most places outside the US and Europe. This lack of ground truth on severe hail puts society and economies at risk. Hail is generated within storms by strong updrafts. These updrafts exhibit unique signatures in NASA and other agency satellite observations, offering new opportunities for hailstorm analysis. Geostationary (GEO) visible and infrared imagery has been collected for ~15-25 years across the world (region dependent) and methods have been developed at NASA Langley Research Center (LaRC) to detect hailstorm updrafts using GEO imagery. Climatological GEO updraft data has been used by Willis Towers Watson (WTW), a leader in catastrophe risk assessment for the insurance industry, and Karlsruhe Institute of Technology to develop CatModels over Europe and Australia. Hail can also be inferred with passive microwave imagery collected by low-Earth-orbiting sensors such as the GPM GMI, TRMM TMI, AMSR-E, AMSR-2, SSM/I, and SSMIS over the last 20+ years using methods developed at the Marshall Space Flight Center (MSFC). Hailstorms generate enhanced lightning flash rates that can be tracked using new GOES-R series GEO Lightning Mapping (GLM) imagery. Atmospheric reanalyses can be used to define favorable hailstorm environments for combination with the satellite-based storm detections. This presentation will describe a framework for developing continental to global hail climatologies and CatModels based on NASA satellite data and capabilities. This is a collaboration between LaRC and MSFC, WTW, and partners in Brazil, Argentina, and South Africa. This project seeks to mitigate hail disasters by aiding development of new satellite-based severe storm nowcasting tools by regional partners and developing climatologies to improve societal understanding of hail frequency. GEOO visible and infrared metrics of storm intensity, environmental conditions based on reanalyses, spotter hail reports and radar MESH observations are intercompared to quantify the detectability of hailstorms, and our ability to discriminate hailstorms from other severe storms. We are also maturing methods using land surface imaging satellite data (e.g. MODIS, Landsat, Sentinel 1 and 2) to identify hail damage to agriculture. Work with WTW will improve socioeconomic resilience through development of new CatModels. Southern Brazil, Uruguay, Paraguay, and Argentina feature some of the most intense thunderstorms on Earth. South America and South Africa are developing insurance markets of interest to WTW clients, and is similar to other regions routinely impacted by hail that do not have comprehensive hail reporting or radars to assess hailstorm frequency. Project datasets will be made available via online GIS-enabled tools developed at the LaRC Atmospheric Science Data Center (ASDC) which will visualize data and provide it in multiple formats for use in a wide range of open source and commercial tools.

Kristopher Michael Bedka↗

MicroBooNE Public Data Sets: a Collaborative Tool for LArTPC Software Development

Among liquid argon time projection chamber (LArTPC) experiments MicroBooNE is the one that continually took physics data for the longest time (2015-2021), and represents the state of the art for reconstruction and analysis with this detector. Recently published analyses include oscillation physics results, searches for anomalies and other BSM signatures, and cross section measurements. LArTPC detectors are being used in current experiments such as ICARUS and SBND, and being planned for future experiments such as DUNE. MicroBooNE has recently released to the public two of its data sets, with the goal of enabling collaborative software developments with other LArTPC experiments and with AI or computing experts. These data sets simulate neutrino interactions on top of off-beam data, which include cosmic ray background and noise. The data sets are released in two formats: the native art/ROOT format used internally by the collaboration and familiar to other LArTPC experts, and the HDF5 format which contains reduced and simplified content and is suitable for usage by the broader community. This contribution presents the open data sets, discusses their motivation, the technical implementation, and the extensive documentation -- all inspired by FAIR principles. Finally, opportunities for collaborations are discussed.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Automated Energy-Dispersive X-ray Spectroscopy Analysis for Multi-Modal Few-Shot Learning

Scanning transmission electron microscopy (STEM) is a powerful tool that allows for the atomic-scale analysis of a materials’ structure, chemistry, and defect domains (Akers et al. 2021). The current generation of microscopes generate vast amounts of data, surpassing the limits of effective manual analysis traditionally performed by domain experts (Spurgeon et al. 2021). While recent strides in machine learning have significantly enhanced the processing of large and intricate datasets acquired through electron microscopy, the prevalent use of proprietary software packages for initial data collection poses a challenge. In many cases, these software packages act as a ‘black box’, constraining user functionality and hindering the output of data in a format that is conducive to seamless integration into machine learning models. This work addresses these challenges by adapting HyperSpy, an open-source Python library, for the analysis and quantification of raw energy dispersive spectroscopy (EDS) data acquired through STEM. The modified HyperSpy code successfully facilitates user-defined segmentation of the data, enabling the integration of atomic %, weight %, and raw EDS spectra for each segmented region into an existing few-shot machine learning model. While initial results reveal discrepancies in quantified atomic and weight percentages when compared to proprietary software, ongoing efforts aim to rectify this issue by refining the fit of the HyperSpy model to the EDS spectra. Overall, this research underscores the potential of open-source tools like HyperSpy to enhance the accessibility of analytical tools, fostering a transparent and user-friendly environment for seamlessly incorporating electron microscopy data into machine learning models.

36 MATERIALS SCIENCE↗

Datum: A Scientific Metadata Catalog

The data catalog market is currently flooded with a myriad of different products, but none serve the scientific community well. There are cloud-native tools like Databricks, Snowflake,to on-premise solutions like Collibra and Datahub. The common failing of all these tools however, is their inability to serve the scientific data community directly. Most catalogs are targeted towards financial, health, or user data - not sensor or scientific domain data. They also prioritize integrations that often don’t exist or are just starting to be used in the scientific realm - all while ignoring common scientific tools and file types. Datum is a catalog which targets the scientific data directly, including the tools and networks in which those tools are used. We work with the producers and consumers of the data where they are, targeting cloud and on-premise with a focus on classified networks. Datum is an Erlang/Elixir application. Technical Features Note: The features listed below are still under development and may change, slightly, upon final delivery of the product. File Formats - Datum has the ability to read additional metadata and provides processing pipelines for the following file formats: Plain Text, PDF, LaTeX, HTML, Open Document Format (.odt), XML, CSV/TSV (and other standard delimiters), OpenDocument Database and Spreadsheets, Geo-Referenced TIFF, Common Data Format, HDF/HDF5, LabView TDMS, Excel, DeltaTables, Parquet, Apache Iceberg, Apache Hudi and many others. Metadata Collection - Scanners for the local and networked file systems and cloud storage providers. Network integration with common databases such as MSSQL and MySQL. User Plugin System - Users are able to provide either file processing, metadata extraction, or sampling plugins in the programming language of their choice. Authentication/Authorization -: OIDC integration, SCIM provisioning and EntraID integration out of the box. Full user and group management system with a “least privilege” operating mode. Governance - Customizable data governance platform; dictate and enforce required metadata, enforce data embargos, and enforce user agreements and NDAs before data access. Ability to create health checks on data, rejecting abandoned or poorly curated data and automatically removing it from the search index. Ability for users to submit corrections. Search - Semantic search is a first class citizen. No licenses to expensive, external software required. Integrated use of vectors and vector-based search allows for AI agent integration at all levels of operation. Metadata Model - Display and control data’s lineage and connections to other data and data directories. Data is modeled after a filesystem - an organization instantly recognizable and navigable by most any user. CLI and SDK - Ships with a Command Line Interface (CLI) tool and with a fully-featured Python SDK. This allows for rapid and programmatic use of Datum by every level of user. Minimal Infrastructure - Datum ships as a single executable file and can be run on any operating system and most CPU architectures. Datum has no reliance on external databases, search indexing tools, or other outside services - and it runs equally well on edge computing devices, cloud services, or in a clustered HPC environment.

darrington, john↗

Gross and Net Soil Methane Flux and Ancillary Data, Edgewater, MD, USA, summer 2022

This data package contains measurements used to quantify methane cycling and environmental conditions in coastal forest soils during the 2022 growing season. It includes time‑series data of soil methane flux, soil respiration, soil temperature, and volumetric water content collected from soil monoliths transplanted along an inundation and salinity gradient. The package also provides one‑time measurements from a stable‑isotope pool‑dilution incubation, including gravimetric water content, methane headspace concentrations, and ¹³CH₄ enrichment over time. Data files are provided in comma‑separated values (CSV) format, with accompanying metadata and readme documentation in PDF and plain‑text formats. All files can be opened with standard software such as R, Python, or spreadsheet programs capable of handling CSV files. The metadata file describes variable definitions, units, processing steps, and the structure of each data table to support reuse and integration with other datasets.

13-C↗