Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Tabular data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Individual Motorist Data - Ohio EV Ownership Trends

The individual motorist dataset contains data and analysis of consumer electric vehicle (EV) ownership trends in rural Appalachian Ohio in comparison with statewide trends. The data span four years, from Q1 2020 to Q2 2023 (partial). They are sourced from the Ohio Bureau of Motor Vehicles registration records and contain detail on drivetrain type (battery-electric vehicle [BEV] or plug-in hybrid electric vehicle [PHEV]); specific vehicle make and model; and registration location at county, city, and ZIP code levels of spatial resolution. Registration data are analyzed at the county level against such indicators as median income, poverty status, urban-rural status, and density of public charging infrastructure. In addition to tabular data, a GIS shapefile with many analysis fields joined is included.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Feature Extraction: Improving Remote Sensor Classification of Non-Proliferation

This research focuses on developing algorithms for nuclear non-proliferation detection using remote sensor modeling. To improve the performance of classification models, we implemented a data pipeline with feature extraction. This pipeline takes raw data and transforms it into smaller data points called features that still describe the model. Improving this classification works towards the departments of energy’s missions of ensuring American’s security and prosperity by creating technology that addresses nuclear challenges. To conduct this analysis, we used the Python programming language and some key packages, including tsfresh and TSFEL. Originally tsfresh was selected because it has the most statistical features out of all the packages. Later TSFEL was incorporated due to the additional features it can extract from data, such as temporal and spectral. However, feature extraction becomes challenging in the presence of missing values. In this case, two additional Python packages were added to our workflow, NumPy and pandas, allowing for the feature extraction process to handle unknown values. Our data pipeline was tested on data collected from a simulation that describes the process state of a physical example. The results show the pipeline’s capability to consume and extract a total 17 features from tabular data. Future work includes producing classifications using decision tree-based models such as XGBoost and improving data collection by analyzing feature importance.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Examining the Characteristics of the Cropland Data Layer in the Context of Estimating Land Cover Change

The United States Department of Agriculture (USDA) Cropland Data Layer (CDL) provides spatially explicit information about crop production area and has served as a prevalent data source for characterizing cropland change in the U.S. in the last decade. Understanding the accuracy of the CDL is paramount because of the reliance on it for management and policy making. This study examined the characteristics of the CDL from 2007 to 2017 using comparisons to other USDA datasets. The results showed when examining the cropland area for the same year, the CDL produced comparable trends with other datasets (R 2 > 0.95), but absolute area differed. The estimated area of cropland changes from 2007 to 2012, 2008 to 2012 and 2012 to 2017 varied from weak to moderate correlation between the CDL and the tabular data (R 2 = 0.005~0.63). Differences in area of cropland change varied widely between data sources with the CDL estimating much larger change area. A series of image processing techniques designed to improve the confidence in cropland change estimated using the CDL reduced the area of estimated cropland change. The techniques also, unexpectedly, lowered the correlation in change estimated between the CDL and the tabular datasets. Estimated land cover change area varied widely based on analyses applied and could reverse from increasing to declining area in cropland. Further analyses showed unlikely change scenarios when comparing different year combinations. The authors recommend the CDL only be used for land cover change analysis if the error can be estimated and is within change estimates.

58 GEOSCIENCES↗

Scaling SQL to the Supercomputer for Interactive Analysis of Simulation Data

AI and simulation workloads consume and generate large amounts of data that need to be searched, transformed and merged with other data. With the goal of treating data as a first-class citizen inside a traditionally compute-centric HPC environment, we explore how the use of accelerators and high-speed interconnects can speed up tasks which otherwise constitute bottlenecks in computational discovery workflows. BlazingSQL is SQL engine that runs natively on NVIDIA GPUs and supports internode communication for fast analytics on terabyte-scale tabular data sets. We show how a fast interconnect improves query performance if leveraged through the Unified Communication X (UCX) middleware. We envision that future computing platforms will integrate accelerated database query capabilities for immediate and interactive analysis of large simulation data.

Glaser, Jens↗

Development, Verification, and Validation of an OpenFOAM-Based Solver for Modeling Inertial Fusion Energy Chambers

Our work seeks to introduce a computational tool tailored to the physics of inertial fusion energy chambers, in particular, those concepts based on thick liquid walls. In this approach, the structural materials are protected by several neutron mean-free-paths of renewable liquid and thus will be able to survive much longer than un-shielded walls, with virtually all structures lasting for the life of the plant and enabling the use of commercially available and qualified materials. The OpenFOAM-based solver named rhoCentralFoam has been used as a starting point. rhoCentralFoam belongs to the standard OpenFOAM solver toolset. It is a high-speed, explicit compressible flow solver with shock-capturing capability. While the main features have been retained, the solver had to be restructured to make use of tabular data for equations of states, a necessary addition to model the complex thermo-physical properties of ionized gasses. This entailed the need to change the independent state variables used by the solver, resulting in a new thermodynamic library and slightly different solution algorithm. Moreover, a radiation heat transfer model based on the P-1 approximation was added to the solver. The solver is verified against an analytical solution from the Sedov-Taylor-Neumann test problem to showcase the ability of the hydrodynamic solvers to handle strong shocks, whereas the P-1 model was verified using a simple one-dimensional problem with an analytical solution. Additionally, a validation case involving shock-wave propagation through jet array is presented, and the results are compared with experimental data from the open literature. Lastly, in order to showcase the utility of the solver for practical cases, we applied the refined solver to two representative scenarios: gas venting within the HYLIFE-II chamber and the compression of the gas following the partial ablation of the liquid wall.

Chamber dynamics↗

Tree-based algorithms for weakly supervised anomaly detection

Weakly supervised methods have emerged as a powerful tool for model-agnostic anomaly detection at the Large Hadron Collider (LHC). While these methods have shown remarkable performance on specific signatures such as dijet resonances, their application in a more model-agnostic manner requires dealing with a larger number of potentially noisy input features. In this paper, we show that using boosted decision trees as classifiers in weakly supervised anomaly detection gives superior performance compared to deep neural networks. Boosted decision trees are well known for their effectiveness in tabular data analysis. Our results show that they not only offer significantly faster training and evaluation times, but they are also robust to a large number of noisy input features. By using advanced gradient boosted decision trees in combination with ensembling techniques and an extended set of features, we significantly improve the performance of weakly supervised methods for anomaly detection at the LHC. This advance is a crucial step toward a more model-agnostic search for new physics. Published by the American Physical Society 2024

Astronomy & Astrophysics↗

Jester

Jester is a Rust CLI designed to package and send time series or tabular data to the data warehouse DeepLynx. It primarily reads .csv files and then sends those via HTTP or Websocket requests to an external instance of DeepLynx. It is meant to run on a host computer which has access to the data.

Darrington, John↗

Deeplynx-loader

'deeplynx-timeseries-loader' is a library designed to make it as easy as possible for users to download and access timeseries or tabular data from DeepLynx.

Darrington, John↗

Python wrapper library and analysis functions for Geotab Altitude API [SWR-24-77]

This software library serves as a Python wrapper for Geotab's Altitude API. It streamlines querying of the API, converts loosely structured API outputs into a standardized tabular data format, and enables analysis of the resulting data tables. It also includes example notebooks showing how to use the library.

Bruchon, Matthew↗

Spark Channel Dynamics of Electrostatic Discharges

When two differently-charged objects are brought in close proximity to each other, the resulting high electric fields can cause electron avalanche breakdown of the air gap separating the objects, a process known as electrostatic discharge (ESD). If enough initial charge is stored on the objects, the electrical breakdown can proceed to ionize the air to such a degree that a highly conductive filament of plasma forms in the gap, known as a spark channel. The spark electrically bridges the air gap, resulting in a rapid pulse of current that neutralizes the charge difference. The current pulse produces significant heating of the gas in the spark, resulting in dissociation, ionization, thermal radiation, and hydrodynamic expansion. ESD presents a hazard to electrically-sensitive devices, with consequences such as economic losses (e.g. damaged electronics) or unsafe response (e.g. unintended ignition of flammable gas mixtures, initiation of detonators, etc.). For this thesis, the ESD spark is taken to occur between two conducting electrodes, with the spark channel being axisymmetric in a cylindrical coordinate system centered on the channel. An RLC-type circuit is used for the discharge model of the ESD event. The spark is treated as a time-dependent resistance that is in series with a capacitance, an inductance, and (optionally) a load resistance representing a “victim” component under threat from the ESD event. The primary motivation of this work is to use a numerical hydrodynamic model to understand the energy dissipation and transport processes in the spark. The model consists of the compressible Euler equations of mass, momentum, and energy conservation together with an Eddington/P1 approximation for thermal radiation transport. To close the hydrodynamic system, an equation of state (EOS) was fitted from tabular data for air that accounts for the dissociation and ionization of air species. The hydrodynamic equations are solved using a conservative Lagrangian finite volume method. These partial differential equations are coupled to the circuit equations by calculation of the spark resistance via numerical integration of the electrical conductivity of the channel. Computational results are compared against experimental measurements of discharge current and radial density of the spark channel.

42 ENGINEERING↗

Probabilistic Forward Modeling of Galaxy Catalogs with Normalizing Flows

Abstract Evaluating the accuracy and calibration of the redshift posteriors produced by photometric redshift (photo- z ) estimators is vital for enabling precision cosmology and extragalactic astrophysics with modern wide-field photometric surveys. Evaluating photo- z posteriors on a per-galaxy basis is difficult, however, as real galaxies have a true redshift but not a true redshift posterior. We introduce PZFlow, a Python package for the probabilistic forward modeling of galaxy catalogs with normalizing flows. For catalogs simulated with PZFlow, there is a natural notion of “true” redshift posteriors that can be used for photo- z validation. We use PZFlow to simulate a photometric galaxy catalog where each galaxy has a redshift, noisy photometry, shape information, and a true redshift posterior. We also demonstrate the use of an ensemble of normalizing flows for photo- z estimation. We discuss how PZFlow will be used to validate the photo- z estimation pipeline of the Dark Energy Science Collaboration, and the wider applicability of PZFlow for statistical modeling of any tabular data.

Astronomy & Astrophysics↗

Oxidation Penetration in Nuclear Graphite

Study results are presented for seven grades of graphite where 21 cylindrical samples were oxidized to a nominal level of mass loss. Low mass loss samples exhibited ~3% oxidative mass loss, intermediate ~8%, and high ~11%. Flat sample surfaces were covered during oxidation to minimize oxidation penetration at the top and bottom of each sample. Oxidized samples had their diameters reduced stepwise in 1 mm or 2 mm increments. Residual samples were weighed, geometric dimensions were recorded, and Archimedes measurements were taken at each step. Preliminary comparative graphical analysis is presented to illustrate the resultant density gradients observed. Raw tabular data are also provided.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Flaw Tolerance Assessment for DOE Standard SNF Dry Storage Canisters - 26550

The U.S. DOE has designed four spent nuclear fuel (SNF) dry storage canisters for storing DOE standardized SNFs. The DOE standard canisters are cylindrical shells with a diameter of 24 inches (610 m) or 18 inches (457 m), a wall thickness of 0.5 inches (12.7 m) or 0.375 inches (9.53 m), and a length of 15 feet (4.57 m) or 10 feet (3.05 m). These DOE canister geometries are completely different from commercial canisters. The latter may experience chloride-induced stress cracking corrosion (CI-SCC) because they are stored near coastal regions. The former may not experience CI-SCC but face different challenges because they are stored in the SNF storage facilities. Because of large residual stresses, mechanical flaws may occur in the DOE canisters during long-distance transportation or lifting handling. To date, only limited structural integrity analyses were carried out through drop tests on the DOE canisters, but a more general flaw tolerance assessment has not been performed. Therefore, the failure assessment diagram (FAD)-based fracture mechanics method, as codified by the latest API 579-1/ASME FFS-1-2021 Edition, is adopted in this work to assess surface flaw tolerance for DOE canisters under operation loading and welding residual stresses (WRS), where the new code-recommended WRS distributions are used. To more adequately consider the transverse distribution of WRS, an equivalent WRS distribution is proposed to account for the WRS reduction with distance from the weld centerline. Moreover, the closed-form solutions of stress intensity factor K, which serves as the crack driving force during subcritical crack growth, are developed from the tabular data of the K factors provided in API 579-1/ASME FFS-1 and used to determine more accurate flaw sizes at flaw instability. Subsequently, the Level 2 assessment procedures with 12 assessment steps, as codified and detailed in API 579-1 and ASME FFS-1, are followed to assess the flaw tolerance for the surface flaws in the DOE standard canisters with consideration of normal or accident operation loads combined with WRS. The assessment results show that the four designs of DOE standard canisters can tolerate all surface flaws that meet the code permitted maximum sizes of a flaw length of 8 inches (i.e., 200 mm) and a flaw depth of 80% wall thickness. This demonstrates that all designs of DOE standard canisters are robust and reliable.

DOE standard canister↗

Meteoric 10Be Flux Calibration Data for the East River Watershed, Colorado, USA

This data package contains tabular and geospatial data used to quantify and model meteoric beryllium-10 fluxes in the East River watershed, Colorado, USA. The tabular component includes calibration-site data from five glacial moraine sites and includes environmental variables used to evaluate spatial controls on meteoric 10Be delivery, including elevation, mean annual precipitation (MAP), mean snow depth, and mean snow water equivalent (SWE). These site-level data were used to compare observed fluxes with environmental gradients across the watershed and to evaluate the effects of erosion correction on flux estimates. The package also includes supporting slope and curvature values used to assess topographic inputs to the erosion analysis. A second component of the data package contains updated manuscript tables and regression outputs used to summarize the relationships between meteoric 10Be flux and environmental predictors. These tables include meteoric 10Be sample information and AMS results, site-level environmental values, site-level meteoric 10Be inventory and flux values, watershed-averaged predicted fluxes, soil bulk density measurements, fine-fraction values, soil pH measurements, and regression statistics including slope, intercept, coefficient of determination, and p-value. The regression products include both standard linear regressions and regressions in which the intercept is constrained to pass through zero, and they support the analyses presented in the companion manuscript. Together, these tabular files provide the numerical basis for the manuscript tables and the regression-based interpretation of meteoric 10Be flux variability in a snow-dominated mountain watershed. The geospatial component of the package consists of GeoTIFF raster files used to generate the map products presented in Figures 2 and 6 of the companion manuscript. These rasters represent watershed-scale spatial layers for environmental variables and regression-based predictions of meteoric 10Be flux. This dataset contains comma-separated values files (.csv), Microsoft Excel files (.xlsx), GeoTIFF raster files (.tif), and upporting metadata files, including CSV data dictionaries and readme text files (.csv, .txt). The tabular files can be opened with standard spreadsheet software, and the raster files can be viewed and analyzed in GIS software such as ArcGIS Pro or QGIS. Together, these files document the numerical and spatial datasets used to calibrate and predict meteoric 10Be delivery in the East River watershed.

East River↗

Conditional distribution estimation of building characteristics with diffusion models for urban energy modeling

Understanding current energy consumption behavior in communities is critical for informing future energy use decisions and enabling efficient energy management. Urban energy models, which are used to simulate these energy use patterns, require large datasets with detailed building characteristics for accurate outcomes. However, such detailed characteristics at the individual building level are often unknown and costly to acquire, or unavailable. Through this work, we propose using a generative modeling approach to generate realistic building attributes to fill in the data gaps and finally provide complete characteristics as inputs to energy models. Our model learns complex, building-level patterns from training on a large-scale residential building stock model containing 2.2 million buildings. We employ a tabular diffusion-based framework that is designed to handle heterogeneous (discrete and continuous) features in tabular building data, such as occupancy, floor area, heating, cooling, and other equipment details. We develop a capability for conditional diffusion, enabling the imputation of missing building characteristics conditioned on known attributes. We conduct a comprehensive validation of our conditional diffusion model, firstly by comparing the generated conditional distributions against the underlying data distribution, and secondly, by performing a case study for a Baltimore residential region, showing the practical utility of our approach. Our work is one of the first to demonstrate the potential of generative modeling to accelerate building energy modeling workflows.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Reducing Memory Consumption in Calico with Shared Memory

This document details the work to reduce memory consumption in Calico. Calico is SimTools’ Constructive Solid Geometry (CSG) and geometry painting library. It is primarily used to paint material volume fractions in the Eulerian meshes of the physics codes. Calico provides point-in-body checks for the geometry supplied by an Oso model, which are then aggregated by the host codes. In addition, Calico can be used to build Oso models and is used by Ingen for that purpose. Oso models, and thus Calico, provide support for various CSG primitives such as spheres, cylinders, surfaces generated by rotating tabular curve data, and STL files as well as binary combinations of those primitives. Prior to refactoring Calico will run out of memory on CTS-1 machines when 36 MPI ranks are used per node when reading STL models on the order of 1.5 GB. This limitation is a bottleneck in designer workflow. This problem has been alleviated through the use of data structures to both reduce memory consumption and to leverage MPI-3 shared memory. This report details the data structures targeted for refactoring in Calico, the methods and implementation details for reducing memory consumption and leveraging shared memory, and results for one test problem. Results show a memory reduction when loading a 1 GB STL file by a factor of 27.5, from 93.4 to 3.4 GB.

97 MATHEMATICS AND COMPUTING↗

GLBRC Soil Yearlong Incubation 13C-SIP-Lipidomics

Data package for Lipids represent a dynamic, yet stable pool of microbially-derived soil carbon This data is published under a CC0 license. The authors encourage data reuse and request attribution by referencing the below citations for the data packages and associated manuscript. Please cite as: Rempfert KR, Bell SL, Kasanke CP, Kyle JE, Hofmockel KS. 2025. GLBRC Soil Yearlong Incubation 13C-SIP-Lipidomics. [Data Set] PNNL DataHub. doi: Rempfert KR, Bell SL, Kasanke CP, Kyle JE, Hofmockel KS. 2025. MSV000097435: GLBRC soil yearlong incubation 13C-SIP-Lipidomics [Data Set] MassIVE. doi:10.25345/C57659T3K Rempfert KR, Bell SL, Kasanke CP, Kyle JE, Hofmockel KS. 2025. Lipids represent a dynamic, yet stable pool of microbially-derived soil carbon. In Prep This data package consists of compound-specific 13C SIP-lipidomics data from a yearlong tracer incubation experiment designed to investigate microbial lipid persistence in switchgrass bioenergy crop soils. In order to explore how lipid structure may modulate the persistence of C in soil lipids, we leveraged soils from two sites (Michigan - sandy texture, Wisconsin - silty texture) operated by the U.S. Department of Energy-funded Great Lakes Bioenergy Research Center (GLBRC). These sites had comparable climates, identical management practices, but contrasting soil textures, allowing us to assess the variability of lipid accrual or degradation in soils as well as provide insight regarding the degree to which edaphic properties may regulate the retention of soil lipids. Untargeted lipidomics analyses were performed to identify 13C-labeled lipids in the soil microbiome after long-term incubation. Soils were supplemented with 100 micrograms glucose per gram dry soil (99 atom % 13C or natural abundance for paired control) and incubated; samples were collected two months and one year after glucose addition. Lipid extracts (MPLEx) were analyzed by LC-MS/MS and identified using LIQUID. Calculation of isotopic enrichment of lipids was performed by targeted approach using TarMet to quantify lipid isotopologues and IsoCorrectoR to correct for natural abundance isotopes. Contents: Data package contents reported here are the first version and contain downstream analysis files for the raw LC-MS mass spectrometry files (.mzXML) deposited at the MassIVE database repository under accession MSV000097435 (80 experimental runs; 5.85 GB) | MassIVE DOI: 10.25345/C57659T3K. Support files include the additional data download 'Read Me' file containing data descriptor information. Reported data download contents are structured for compliance with project data sharing guidelines, community standards initiatives, and sponsor stakeholder policies supporting FAIR data principles. Data processing software, analysis tools, and data workflows are listed below corresponding to the host repository long-term location. Available Data Downloads (0.3 GB): "GLBRC soil yearlong incubation 13C-SIP-Lipidomics_readme.txt" - 'Read Me' data package content file (txt) "GLBRC_DataPackage_analysis files" - Data processing files (Rmd) and saved intermediate data processing outputs (rds, csv, xlsx) "GLBRC_13C_lipidomics_dataset.xlsx" - processed data in tabular format (xlsx) Linked Software: LIQUID LC-MS Analysis Software | 10.5281/zenodo.6459462 Lipid Mini-On Software Tools | 10.5281/zenodo.1492803 pmartR Omics Statistical Software | 10.5281/zenodo.6108667 xcms (v4.3.3) TarMet (v1.1.1) IsoCorrectoR (1.24.0) Funding Acknowledgments: This research was supported by an Early Career Research Program award funded by the U.S. Department of Energy, Office of Science, Office of Biological and Environmental Research (OBER) Genomic Science program under FWP 68292, FWP 07880 and EMSL Exploratory Research Project 51095. A portion of this work was performed in the William R. Wiley Environmental Molecular Sciences Laboratory, a national scientific user facility sponsored by OBER and located at Pacific Northwest National Laboratory (PNNL). PNNL is a multi-program national laboratory operated by Battelle for the DOE under Contract DE-AC05-76RLO1830.

Rempfert, Kaitlin R [Pacific Northwest National La↗

Data and scripts associated with “Non-random processes impacting organic matter chemistry are maximized in mid-order streams”

NOTE: The manuscript associated with this data package is currently in review. The data may be revised based on reviewer feedback. Upon manuscript acceptance, this data package will be updated with the final dataset and additional metadata. This data package is associated with the publication “Non-random processes impacting organic matter chemistry are maximized in mid-order streams” submitted to Limnology and Oceanography (L&O) by Danczak et al. (in review). This package contains data and scripts used to investigate dissolved organic matter (DOM) molecular chemistry and diversification processes across 47 surface-water sampling sites in the Yakima River Basin, Washington, USA, during an August 2021 sampling campaign. The package contains analyses of ultrahigh-resolution Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS), geochemical measurements, geospatial attributes, molecular diversity, and meta-metabolome ecological null models needed to reproduce the main manuscript results. The underlying field data were pulled from exising data packages at https://doi.org/10.15485/1892052 (Fulton et al., 2022) and https://doi.org/10.15485/1898914 (Grieger et al., 2022). For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. We thank the following organizations for providing access to field locations for sample collection: the United States Forest Service, Washington Department of Fish and Wildlife, Washington Department of Natural Resources, the Confederated Tribes and Bands of the Yakama Nation, and the Cowiche Canyon Conservatory. Research was conducted under Washington State Parks and Recreation Commission Scientific Research Permit #210901. We are grateful to the Yakama Nation Tribal Council and Yakama Nation Fisheries for their collaboration in facilitating sample collection and ensuring data usage aligns with their values and worldview. This data package contains an R-Markdown file for analyses and five folders: (1) Data, (2) Geospatial Data, (3) Supplemental_Files, (5) Figures_pdf, (4) and src. The Data folder contains tabular inputs and derived files used in the manuscript analysis. The Geospatial Data folder contains climate and water-balance, hydrologic, land-cover, population/regional water-use, stream, topographic, and stream-order attribute CSV files. The src folder contains scripts used to process data, run analyses, and generate figures. The Figures_pdf folder contains manuscript figure outputs. The Supplemental_Files folder contains supplemental analysis products. All files are .csv, .pdf, .html, .png, .R, .Rmd, .svg, or .tre. This data package is associated with the rcfsa-RC2-SPS_Null_Modeling repository found at https://github.com/river-corridors-sfa/rcfsa-RC2-SPS_Null_Modeling.

54 ENVIRONMENTAL SCIENCES↗