Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “open data format”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

HopPyBar

HopPyBar is a python program to import, analyze, and export split-Hopkinson pressure bar (SHPB, also known as Kolsky bar) data. Traditional analysis offers a black box approach, where input data is converted to analyzed output by performing a series of calculations without user involvement. This program serves as a developmental platform to "white box" the data analysis process. Data streams can be captured (in-situ) to enable advanced or unconventional analyses, statistics, and comparisons. Additionally, the program is geared towards the standardized forms of input and output used at LANL to streamline analysis, but the open nature of the program makes additional input/output schemes straightforward to add. General workflow will import SHPB data in one of a number of formats, identify relevant portions of data signals, and convert to stress-strain-strain rate to show material behavior as a function of dynamic testing.

Morrow, Benjamin↗

OASIS: Open-Source AI Software Infrastructure for Science-SBIR Phase I

OASIS: Open source AI Software Infrastructure for Science is developed for researchers in the scientific domain. OASIS provides a data API to ingest and serve scientific data formats and annotations within AI workflows. It delivers a unique integration of features such as coupling of data to AI models, scalable training, cloud deployment into a cohesive web and command-line interface, and state-of-the-art techniques to debug and enhance AI models.

Chaudhary, Aashish↗

Conversion Helper 4 Easy Serialization Of Exi (ch4ese)

CH4ESE is an EXI conversion tool developed in Python3 that utilizes the open-source EXIficient implementation of the W3C EXI format specification. CH4ESE can be used to translate to and from EXI format using the command line with input data or using the web server for live-translation.

Rohde, KennethW [Idaho National Laboratory (INL), ↗

Globular Cluster Candidates in the Sagittarius Dwarf Galaxy

Recently, new Sagittarius (Sgr) dwarf-galaxy globular clusters were discovered, which opens the question of the actual size of the Sgr globular cluster population, and therefore on our understanding of the Sgr galaxy formation and accretion history of the Milky Way. Based on Gaia EDR3 and SDSS IV DR16 (APOGEE-2) data sets, we performed an analysis of the color–magnitude diagrams (CMDs) of the eight new Sgr globular clusters found by Minniti et al. from a sound cleaning of the contamination of Milky Way and Sgr field stars, complemented by available kinematic and metal abundance information. The cleaned CMDs and spatial stellar distibutions reveal the presence of stars with a wide range of cluster membership probabilities. Minni 332 turned out to be a younger (<9 Gyr) and more metal-rich ([M/H] ≳ -1.0 dex) globular cluster than M54, the nuclear Sgr globular cluster; as could also be the case of Minni 342, 348, and 349, although their results are less convincing. Minni 341 could be an open cluster candidate (age < 1 Gyr, [M/H] ~ -0.3 dex), while the analyses of Minni 335, 343, and 344 did not allow us to confirm their physical reality. We also built the Sgr cluster frequency (CF) using available ages of the Sgr globular clusters and compared it with that obtained from the Sgr star formation history. Both CFs are in excellent agreement. However, the addition of eight new globular clusters with ages and metallicities distributed according to the Sgr age–metallicity relationship turns out in a remarkably different CF.

79 ASTRONOMY AND ASTROPHYSICS↗

A Machine Learning Framework to Deconstruct the Primary Drivers for Electricity Market Price Events

As the electricity grid is moving towards a 100% Renewable Energy Source Bulk Power Grid, the overall operations of the power system operations and electricity markets are changing. The electricity markets are not only dispatching resources economically but also taking into account various controllable actions like renewable curtailment, transmission congestion mitigation, and energy storage optimization to make sure the grid is operating reliably. As a result, price formations in electricity markets have become quite complex. Traditional root cause analysis and statistical approaches are rendered inapplicable to analyze and infer the main drivers behind price formation in the modern grid and markets with variable renewable energy (VRE). In this paper, we propose a machine learning analysis framework to deconstruct some primary drivers for price formation in modern electricity markets with high renewable energy and the outcomes can be utilized for various critical aspects of market design, renewable dispatch and curtailment, operations, and cyber-security applications. The framework can be applied to any ISO or market data and in this paper it is applied to open-source publicly available datasets from California Independent System Operator (CAISO) and ISO New England.

machine learning (ML), electricity markets, Renewa↗

A Simple Standard for Sharing Ontological Mappings (SSSOM)

Abstract Despite progress in the development of standards for describing and exchanging scientific information, the lack of easy-to-use standards for mapping between different representations of the same or similar objects in different databases poses a major impediment to data integration and interoperability. Mappings often lack the metadata needed to be correctly interpreted and applied. For example, are two terms equivalent or merely related? Are they narrow or broad matches? Or are they associated in some other way? Such relationships between the mapped terms are often not documented, which leads to incorrect assumptions and makes them hard to use in scenarios that require a high degree of precision (such as diagnostics or risk prediction). Furthermore, the lack of descriptions of how mappings were done makes it hard to combine and reconcile mappings, particularly curated and automated ones. We have developed the Simple Standard for Sharing Ontological Mappings (SSSOM) which addresses these problems by: (i) Introducing a machine-readable and extensible vocabulary to describe metadata that makes imprecision, inaccuracy and incompleteness in mappings explicit. (ii) Defining an easy-to-use simple table-based format that can be integrated into existing data science pipelines without the need to parse or query ontologies, and that integrates seamlessly with Linked Data principles. (iii) Implementing open and community-driven collaborative workflows that are designed to evolve the standard continuously to address changing requirements and mapping practices. (iv) Providing reference tools and software libraries for working with the standard. In this paper, we present the SSSOM standard, describe several use cases in detail and survey some of the existing work on standardizing the exchange of mappings, with the goal of making mappings Findable, Accessible, Interoperable and Reusable (FAIR). The SSSOM specification can be found at http://w3id.org/sssom/spec. Database URL: http://w3id.org/sssom/spec

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

pvlib iotools—Open-source Python functions for seamless access to solar irradiance data

Access to accurate solar resource data is critical for numerous applications, including estimating the yield of solar energy systems, developing radiation models, and validating irradiance datasets. However, lack of standardization in data formats and access interfaces across providers constitutes a major barrier to entry for new users. pvlib python’s iotools subpackage aims to solve this issue by providing standardized Python functions for reading local files and retrieving data from external providers. All functions follow a uniform pattern and return convenient data outputs, allowing users to seamlessly switch between data providers and explore alternative datasets. The pvlib package is community-developed on GitHub: https://github.com/pvlib/pvlib-python. As of pvlib python version 0.9.5, the iotools subpackage supports 12 different datasets, including ground measurement, reanalysis, and satellite-derived irradiance data. The supported ground measurement networks include the Baseline Surface Radiation Network (BSRN), NREL MIDC, SRML, SOLRAD, SURFRAD, and the US Climate Reference Network (CRN). Additionally, satellite-derived and reanalysis irradiance data from the following sources are supported: PVGIS (SARAH & ERA5), NSRDB PSM3, and CAMS Radiation Service (including McClear clear-sky irradiance).

14 SOLAR ENERGY↗

MolViewSpec: a Mol* extension for describing and sharing molecular visualizations

Data visualization is a pivotal component of a structural biologist’s arsenal. The Mol* Viewer makes molecular visualizations available to broader audiences via most web browsers. While Mol* provides a wide range of functionality, it has a steep learning curve and is only available via a JavaScript interface. To enhance the accessibility and usability of web-based molecular visualization, we introduce MolViewSpec (molstar.org/mol-view-spec), a standardized approach for defining molecular visualizations that decouples the definition of complex molecular scenes from their rendering. Scene definition can include references to commonly used structural, volumetric, and annotation data formats together with a description of how the data should be visualized and paired with optional annotations specifying colors, labels, measurements, and custom 3D geometries. Developed as an open standard, this solution paves the way for broader interoperability and support across different programming languages and molecular viewers, enabling more streamlined, standardized, and reproducible visual molecular analyses. MolViewSpec is freely available as a Mol* extension and a standalone Python package.

Midlik, Adam [European Bioinformatics Institute (U↗

CoCoMET v1.0: a unified open-source toolkit for atmospheric object tracking and analysis

Advances in performance and analysis capabilities have accelerated the development of object tracking algorithms for atmospheric research. This has resulted in a growing number of studies using Lagrangian tracking techniques to analyze the evolution of atmospheric phenomena and the underlying processes. However, the increasing complexity and variety of tracking algorithms present a steep learning curve for new users and make it difficult for existing users to compare algorithm performance. We introduce CoCoMET (Community Cloud Model Evaluation Toolkit), an open-source toolkit that addresses these issues. CoCoMET simplifies the process of running multiple tracking algorithms simultaneously and analyzing objects in both model and observational datasets by specifying parameters in a single configuration file. It standardizes input data from different sources into a consistent format and unifies the tracking output across algorithms. CoCoMET enhances the functionality of existing tracking methods by calculating additional properties such as cell growth and dissipation rates, perimeter, surface area, convexity, and irregularity. In addition, CoCoMET includes a novel method for identifying mergers and splits in 2D and 3D tracks and supports the integration of Eulerian/stationary datasets external to the tracking data for process studies. Its potential utility is demonstrated through examples of model intercomparison, model evaluation against observations, and comparisons between tracking algorithms. Designed for open-source environments, CoCoMET will continue to expand with future releases, incorporating more input data types and tracking algorithms.

54 ENVIRONMENTAL SCIENCES↗

Catalytic hydrogenation of HMF to BHMF over copper catalysts

2,5-Bis(hydroxymethyl)furan (BHMF) is a bio-derived building block for polyester production, obtained via the hydrogenation of 5-hydroxymethylfurfural (HMF). First-principles thermodynamic equilibrium calculations indicate that this reaction is not thermodynamically limited under relevant conditions (e.g., 100 °C and high H 2 partial pressure). In this work, crude HMF was employed as the feedstock for BHMF synthesis. Initially, acidic impurities and humins were removed from unrefined HMF through filtration using a packed bed of γ-alumina. A comprehensive study of the filtration process is presented, including filtration kinetics, breakthrough curve analysis, and mathematical modeling. The purified HMF was subsequently hydrogenated over a 10 wt% CuZrO 2 catalyst, using ethanol as the reaction solvent. Batch reactions were first performed for collection of kinetic data to guide the transition to continuous flow operation. Kinetic data was collected in a fixed bed reactor at varying contact time, time on stream, temperature, and HMF concentration. This data was used to develop a kinetic model for HMF hydrogenation. Maximum BHMF production rates were achieved at 130 °C, accompanied by minor formation of byproducts from BHMF ring-opening reactions. The BHMF selectivity was 100 % at 100 °C although with lower reaction rates. Furthermore, catalyst stability tests revealed a loss of up to 50 % in catalytic activity within the first 24 h, likely due to the adsorption of HMF-derived oligomers that are not easily removed by filtration.

Crude HMF filtration↗

COMPASS-FME Terrestrial Ecosystem Manipulation to Probe the Effects of Storm Treatments (TEMPEST) Experiment Level 2 Sensor Data v2-1

This is the version v2-1 Level 2 (L2) data release for COMPASS-FME environmental sensors located at our Terrestrial Ecosystem Manipulation to Probe the Effects of Storm Treatments (TEMPEST) experimental site. This manipulative, ecosystem-scale TEMPEST experiment addresses the potential for freshwater and estuarine-water disturbance events to alter tree function, species composition, and ecosystem processes in a deciduous coastal forest in MD, USA. The experiment uses a large-unit (2000 m2), un-replicated experimental design, with three 50 m × 40 m plots serving as control, freshwater, and estuarine-water treatments. Level 2 (L2) data consist of sensor observations from the COMPASS-FME synoptic sites, TEMPEST, and DELUGE. Compared to the L1 data, these are more consistent (always 15-minute timestamps for the entire year); better QA/QC’d (out of bounds, out of service, and extreme outlier values are removed); and more complete, with a gap-filled time series available alongside the main observations, and additional derived (calculated) variables. L2 data are intended to be rapidly and easily usable in analyses and simulations. However, algorithmic outlier identification always carries the risk of removing valid data, and Level 1 data may be more suitable for analyses that focus on variability or extreme events. This dataset includes: - An overall dataset README file that describes the current version, gives citation and contact information, etc. - Site- and year-specific folders, each holding variable-specific Parquet (a high performance, space efficient format; see https://parquet.apache.org) data files for each site and plot in that year. - Metadata files within each site-year folder provide full information on data units, expected ranges, contact information, detailed flood times, as well as a general description of the site. - Environmental sensor types that appear in the data files include weather (ClimaVUE50, CS, RM Young, and LI instruments in the graphs below); soil conditions (TEROS12); soil redox state (Redox); groundwater variables (AquaTROLL200 and AquaTROLL600); open water sondes (Exo); tree sap velocity (Sapflow); and system voltage and state (Datalogger). Data are reported every 15 minutes. Please see v2-1 TEMPEST L2 Sensor Package Quick Start.pdf for detailed information on data package structure, temporal coverage, and versioning. Data files are in Apache Parquet, a high performance, space efficient format for tabular data. These files can be read using R's `arrow` package (https://arrow.apache.org/docs/r/), with similar tools available in other languages. The TEMPEST flood events occurred on the following dates. They lasted for ~10 hours each day and delivered ~80,000 gallons to each plot; many data streams are available at 1 or 5 minute frequency during these periods. * Tests: Aug 25 (fresh plot) and Sep 9 (salt plot), 2021 * TEMPEST 1: June 22, 2022 * TEMPEST 2: June 6-7, 2023 * TEMPEST 3: June 11-13, 2024

EARTH SCIENCE > ATMOSPHERE > ATMOSPHERIC TEMPERATU↗

2025 TEM Workshop

The TEM Data Management Workshop will take place on August 26 from 9 a.m. to 12 p.m. MT, and will be held virtually on TEAMS. The primary goal of this workshop is to engage NSUF users and stakeholders in discussions about the data needs for the utilization of AI and ML in the analysis of TEM data. Key topics to be covered include data storage, data sharing, data tagging, metadata inclusion, standardized data formats, data augmentation, and annotated training datasets. Additionally, the workshop will provide valuable insights into resources such as the Nuclear Research Data System (NRDS) for data storage and sharing, as well as open-source codes for data analysis.

Bachhav, Mukesh↗

A DECADE of dwarfs: first detection of weak lensing around spectroscopically confirmed low-mass galaxies

We present the first detection of weak gravitational lensing around spectroscopically confirmed dwarf galaxies, using the large overlap between DESI DR1 spectroscopic data and DECADE/DES weak lensing catalogs. A clean dwarf galaxy sample with well-defined redshift and stellar mass cuts enables excess surface mass density measurements in two stellar mass bins ($\log \rm{M}_*=[8.2, 9.2]~M_\odot$ and $\log \rm{M}_*=[9.2, 10.2]~M_\odot$), with signal-to-noise ratios of $5.6$ and $12.4$ respectively. This signal-to-noise drops to $4.5$ and $9.2$ respectively for measurements without applying individual inverse probability (IIP) weights, which mitigates fiber incompleteness from DESI's targeting. The measurements are robust against variations in stellar mass estimates, photometric shredding, and lensing calibration systematics. Using a simulation-based modeling framework with stellar mass function priors, we constrain the stellar mass-halo mass relation and find a satellite fraction of $\simeq 0.3$, which is higher than previous photometric studies but $1.5σ$ lower than $Λ$CDM predictions. We find that IIP weights have a significant impact on lensing measurements and can change the inferred $f_{\rm{sat}}$ by a factor of two, highlighting the need for accurate fiber incompleteness corrections for dwarf galaxy samples. Our results open a new observational window into the galaxy-halo connection at low masses, showing that future massively multiplexed spectroscopic observations and weak lensing data will enable stringent tests of galaxy formation models and $Λ$CDM predictions.

To, Chun-Hao [Chicago U., Astron. Astrophys. Ctr.;↗

A novel approach for adaptive skeleton toolpath generation

Industry 4.0 is revolutionizing manufacturing through the integration of automation and real-time data sharing in cyber-physical systems. At the forefront of this revolution is large-format additive manufacturing. In large-format printing, parts are often designed to be an even number of beads wide to produce a completely dense part. However, voids can still arise. This is often due to the part not being an even number of bead widths wide in some areas, or in geometry containing acute angles, as the process of generating closed contours cannot completely fill the space. Voids can be tolerated in smaller models, but in large-format additive manufacturing they may cause mechanical defects. To fill these voids, open loop paths called skeletons are often used, but they are typically limited by the physical constraints defined in the slicing software. To address this, researchers at Oak Ridge National Laboratory have extended skeleton toolpaths via an adaptive methodology. These adaptive skeletons were found to better fill void spaces through manipulation of physical parameters of the build process and were calculated as part of the slicing process.

42 ENGINEERING↗

Utah FORGE: Powder X-ray Diffraction Data from Well 16A(78)-32 Core

This dataset from Lawrence Livermore National Laboratory (LLNL) consists of four raw X-ray diffraction (XRD) scans and preliminary results of quantitative XRD analysis. The scanned samples were prepared from four subcores, which came from various depths of the FORGE well 16A(78)-32 core. Desired core lengths were selected from available core photos (on GDR), provided by FORGE personnel, and subcored at LLNL. The XRD scans were collected in May 2023 at LLNL as pre-experimental characterization data for these subcores, which will be used in core-flooding experiments at LLNL and in triaxial direct shear experiments at Los Alamos National Laboratory as part of DOE Project 5-2428. XRD scans are in RAW file format (e.g., FORGE-5477-full.raw) and are suitable for viewing and analysis using open-source quantitative XRD software (e.g., Profex; www.profex-xrd.org ) and/or other proprietary instrument software.

15 GEOTHERMAL ENERGY↗

Model Data for the Mesh Convergence Study Demonstrating Benefits of Mixed-polyhedral Mesh in Integrated Hydrology Simulations

This archived model data is related to a study introducing a unique method that employs a stream-aligned mixed-polyhedral mesh to effectively and accurately represent river valleys, stream corridors, and narrow engineered channels in integrated hydrology simulations. The study finds that utilizing stream-aligned mixed-polyhedral meshes in integrated hydrology simulations achieves accuracy on par with a finely refined TIN-based mesh while markedly diminishing computational costs. This archive contains scripts and data files needed to generate the ATS model input, including mesh and ATS input files, for all mesh scenarios using the Watershed Workflow package. Additionally, this archive also provides key outputs from the model simulations that are used in the analysis and post-processing scripts to reproduce figures in the manuscript. The Watershed Workflow package is implemented in Python3. The Jupyter notebooks can be executed through multiple open-source tools, for example, Anaconda Jupyter Lab, VS Studio Code, etc. Other data files include CSV and HDF5 files, which can be read through Python scripts. The input files for the ATS model, open-source integrated hydrology, and transport model are in XML format and can be edited in any commonly used text editors.

54 ENVIRONMENTAL SCIENCES↗

Modeling the Effects of Artificial Drainage on Agriculture-dominated Watersheds using a Fully Distributed Integrated Hydrology Model: Datasets, scripts, model files

This model-data archive supports the research paper that demonstrates the integration of agricultural drainage features—specifically, narrow engineered ditches and tile drains—into a fully distributed, basin-scale integrated surface-subsurface hydrology model (ISSHM), Amanzi-ATS. The model employs innovative computational meshes aligned with agricultural ditches and incorporates the physically based Hooghoudt's drainage equation to simulate tile drainage, offering a novel strategy that enhances the accuracy of hydrological simulations.The archived dataset includes input parameters, model configurations, and select simulation outputs for the Amanzi-ATS model that successfully captured the streamflow patterns in the Portage River Watershed as validated by USGS gauge readings. Jupyter notebook for the preparation of model inputs and post-processing of outputs are also included. The model's predictive performance achieved a normalized Kling-Gupta Efficiency (KGE) of 0.81, surpassing SWAT without the necessity for site-specific calibration.The Amanzi-ATS model presented in this modeL-data archive allows for numerical experiments to explore the shifts in the flow structure under different drainage scenarios. As a tool for advancing the understanding of distributed hydrological responses and nutrient cycling, this archived model provides valuable insights for researchers, modelers, and decision-makers involved in watershed management and environmental modeling.The Watershed Workflow package is implemented in Python3. The Jupyter notebooks can be executed through multiple open-source tools, for example, Anaconda Jupyter Lab, VS Studio Code, etc. Other data files include CSV and HDF5 files, which can be read through Python scripts. The input files for the ATS model, open-source integrated hydrology, and transport model, are in XML format and can be edited in any commonly used text editors.

54 ENVIRONMENTAL SCIENCES↗

The Open Cluster Chemical Abundances and Mapping Survey. VII. APOGEE DR17 [C/N]–Age Calibration

Large-scale surveys open the possibility to investigate Galactic evolution both chemically and kinematically; however, reliable stellar ages remain a major challenge. Detailed chemical information provided by high-resolution spectroscopic surveys of the stars in clusters can be used as a means to calibrate recently developed chemical tools for age-dating field stars. Using data from the Open Cluster Abundances and Mapping survey, based on the Sloan Digital Sky Survey/Apache Point Observatory Galactic Evolution Experiment 2 survey, we derive a new empirical relationship between open cluster stellar ages and the carbon-to-nitrogen ([C/N]) abundance ratios for evolved stars, primarily those on the red giant branch. With this calibration, [C/N] can be used as a chemical clock for evolved field stars to investigate the formation and evolution of different parts of our Galaxy. We explore how mixing effects at different stellar evolutionary phases, like the red clump, affect the derived calibration. We have established the [C/N]–age calibration for APOGEE Data Release 17 (DR17) giant star abundances to be $\mathrm{log}{[\mathrm{Age}(\mathrm{yr})]}_{\mathrm{DR}17}=10.14\,(\pm 0.08)+2.23(\pm 0.19)\,[{\rm{C}}/{\rm{N}}]$, usable for $8.62\leqslant \mathrm{log}(\mathrm{Age}[\mathrm{yr}])\leqslant 9.82$, derived from a uniform sample of 49 clusters observed as part of APOGEE DR17 applicable primarily to metal-rich, thin- and thick-disk giant stars. This measured [C/N]–age APOGEE DR17 calibration is also shown to be consistent with asteroseismic ages derived from Kepler photometry.

79 ASTRONOMY AND ASTROPHYSICS↗