Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Tabular data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Equation-of-state, sound speed, and reshock of shock-compressed fluid carbon dioxide

Mechanical equation-of-state data of initially liquid and solid CO2 shock-compressed to terapascal conditions are reported. Diamond-sapphire anvil cells were used to vary the initial density and state of CO2 samples that were then further compressed with laser-driven shock waves, resulting in a data set from which precise derivative quantities, including Grüneisen parameter and sound speed, are determined. Reshock states are measured to 800 GPa and map the same pressure-density conditions as the single shock using different thermodynamic paths. The compressibility data reported here do not support current density-functional-theory calculations, but are better represented by tabular equation-of-state models.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Pennsylvania Department of Environmental Protection (PA DEP) 26r Detailed Produced Water Compositions (version 1.0)

A database of geochemical compositions of aqueous species in produced water reported to the PA DEP. Samples were collected between mid-2012 to early-2020. Data from publicly-available PA DEP 26r reports were scraped from pdf files and cumulated into tabular spreadsheet format for >1000 produced water streams from Marcellus wells in Pennsylvania. In addition to providing the original values, the NETL NEWTS team has reformatted the dataset to allow sample streams to be easily copied into OLI Studio and Geochemist WorkBench (GWB) software for modeling the geochemistry and the recovery of critical minerals, such as lithium, from these produced water streams. In addition, a version of the dataset has been included with predictions for some missing values in the original dataset using machine learning techniques within CoDaRT software, a public ML software developed by the Nation Energy Technology Laboratory. We have made the Input into CoDaRT and one example output from CoDaRT available in this dataset.

Aqueous Chemistry↗

Model Data Archive Associated with Manuscript "Fire-altered Carbon Pools Create Disturbance Memory in Stream Dissolved Organic Carbon"

This data package supports the publication “Fire-altered Carbon Pools Create Disturbance Memory in Stream Dissolved Organic Carbon” by Li et al. (2026). The package contains processed model inputs, configuration files, restart files, simulation outputs, scripts, and visualization products used to evaluate post-fire dissolved organic carbon (DOC) dynamics in the Naches River Watershed, Washington, USA, following the 2021 Schneider Springs Fire. The modeling workflow couples ELM-BGC, the biogeochemistry-enabled Energy Exascale Earth System Model Land Model; ATS, the Advanced Terrestrial Simulator for integrated surface-subsurface hydrology; and PFLOTRAN, a reactive transport model for multicomponent aqueous geochemistry. Together, these models simulate how wildfire-induced changes in vegetation, litter, coarse woody debris, and soil organic matter influence DOC production, transport, and reaction from burned hillslopes to stream networks. The archive includes preprocessed meteorological, geospatial, hydrologic, and biogeochemical forcing data; ELM-BGC-derived DOC source terms; ATS mesh files; PFLOTRAN reactive-transport inputs; model configuration files; spin-up and transient restart files; watershed-scale diagnostic outputs; stream concentration time series; and figures or visualization files used to inspect and reproduce key results. File types include Hierarchical Data Format 5 (HDF5) files for gridded forcing and model-coupling data, model input and configuration files for ELM-BGC, ATS, and PFLOTRAN, restart and simulation-output files generated by the modeling workflow, tabular or time-series diagnostic outputs, scripts for post-processing and figure generation, and image or visualization products associated with the manuscript. Use of the package depends on the intended task. Re-running the simulations requires the relevant modeling software, including ELM-BGC, ATS, and PFLOTRAN as ATS's geochemical engine. Inspecting outputs and reproducing figures requires Python with scientific plotting libraries such as Matplotlib, and three-dimensional model outputs may be viewed with ParaView. Geographic information system files or maps may be inspected with ArcGIS Pro or comparable GIS software. The data package is intended to enable traceability, reuse, and partial reproduction of the coupled land-to-watershed hydro-biogeochemical modeling workflow used to test how wildfire disturbance affects terrestrial carbon pools and downstream DOC dynamics.

ATS↗

Equation of state of boron carbide B 4 ⁢C

We present the results of recent experiments conducted on the Sandia Z machine and a new tabular equation of state for B 4 ⁢C. The equation of state was calibrated to a combination of density functional calculations reported here and fits to preexisting data. It was constructed partly to recover the effects of a shock-driven, polymorphic phase transition of unknown character beginning at particle velocities of just under 3 km/s (shock pressures of 95 GPa). Some of the Z experiments included sound speeds determined by the overtaking rarefaction method, from which we calculate the Grüneisen parameter and compare with previous experiments conducted at the OMEGA laser [Fratanduono et al ., Phys. Rev. B 94 , 184107 (2016)], our own first principles calculations, and another recent tabular equation of state [Zhang et al ., Phys. Rev. E 102 , 053203 (2020)]. We also compare our results with previous static compression, thermophysical, and melt studies, finding mixed consistency. We predict the onset and completion of shock melting at 225 and 265 GPa, respectively, and predict a melt curve that is largely flat to pressures of several hundred GPa.

36 MATERIALS SCIENCE↗

Efficient Clustering of Software Vulnerabilities using Self Organizing Map (SOM)

The common vulnerabilities and exposures (CVE) database was created with a mission to ``identify, define, and catalog publicly disclosed cybersecurity vulnerabilities''. This rich body of information can be used to enable rapid and efficient response to secure and defend cyber operations and protect critical cyber infrastructure. The main goal of this paper is to develop a visual analytics tool to enable deep analysis of CVEs using unsupervised clustering techniques. We enhance our analysis by first mapping CVEs to hierarchical-classes in Common Weakness Enumeration (CWE) using information in the National Vulnerability Database (NVD). Both the mapping and the numerical representation of CVEs are enabled by V2W-BERT, which uses natural language processing of the extensive information in NVD to generate a large tabular database of 137,226 CVE entries from 1999 to 2020, where each CVE is represented by a vector of 768 numerical features. The vectorized data is processed by Self-Organizing Maps (SOM), which is an unsupervised machine learning technique for dimensionality reduction, visual representation and clustering. Using a Torus map of 6417 units, we achieve ~10-fold data compression of ~140k CVEs using SOM. The trained map is further clustered using standard K-means clustering into 138 clusters of CVEs. We conducted a brief investigation of the rich mapping of CVEs to best-matching-units to K-means clusters, as well as CVEs to CWEs. For example, this novel mapping provided insight into the role of CWE-59 and CWE-264 in several CVEs that is otherwise hard to explore in the original data. We conclude that our this novel approach will not only enable deep analysis of the complex relationships between CVEs and CWEs, but also a mechanism to quickly respond to and design mitigation actions for rapidly evolving vulnerabilities that have not been mapped to existing CWEs.

Panchal, Khyati↗

A Deterministic Multivariate Clustering Method for Drive Cycle Generation from In-Use Vehicle Data

Accurately characterizing vehicle drive cycles plays a fundamental role in assessing the performance of new vehicle technologies. Repeatable, short duration representative drive cycles facilitate more informed decision making, resulting in improved test procedures and more successful vehicle designs. With continued growth in the deployment of onboard telematics systems employing global positioning systems (GPS), large scale, low cost collection of real-world vehicle drive cycle data has become a reality. As a result of these technological advances, researchers, designers, and engineers are no longer constrained by lack of operating data when developing and optimizing technology, but rather by resources available for testing and simulation. Experimental testing is expensive and time consuming, therefore the need exists for a fast and accurate means of generating representative cycles from large volumes of real-world driving data. This paper explores the development and initial validation of a method of generating representative drive cycles from large collections of real-world vehicle data using a deterministic multivariate clustering approach. Starting with theory and diving into the methodology behind representative cycle generation, the paper aims to also present graphical and tabular results of initial validation via vehicle simulation and chassis dynamometer testing. Additional topics for further research and areas for ongoing development will also be presented.

47 OTHER INSTRUMENTATION↗

Geothermal Fault Zone and Fluid Imaging through Joint Airborne ZTEM and Ground MT Data Inversion Analysis

This project has aimed to achieve detailed electrical resistivity resolution at geothermal reservoir scales by combining airborne natural electromagnetic (EM) field surveying (ZTEM) with ground magnetotelluric (MT) measurements to approximate an airborne MT geophysical method. MT alone is relatively expensive and may have permitting challenges in sensitive areas. Airborne ZTEM field data contains only the magnetic field, requires a background assumption, and has been limited to relatively high frequencies, thus suffering uniqueness problems. Based on proto-type 2D simulations, ZTEM ambiguities may be reduced through formal incorporation with possibly sparse ground MT soundings, which we pursued in full 3D for this project. The methodology was tested at the high-temperature Roosevelt Hot Springs geothermal system, Utah, which was considered advantageous given the near total exposure of crystalline reservoir rocks across the project area. ZTEM and ground MT survey data were acquired in 2017, subcontracted to outside parties with which we have worked in the past. These included 80 remote-referenced tensor MT soundings over the Mineral Mountains and adjacent Roosevelt Hot Spring producing geothermal system. These MT stations abut later coverage of a similar number of MT stations taken for the Utah FORGE project providing excellent total data aperture to re-solve structure beneath both project areas better than either set alone. The airborne ZTEM survey covered 704 line kilometers in E-W flight lines with a 250 m line spacing. Although this survey was timed during a maintenance-related shutdown of power production at the Roosevelt Hot Springs, other noise sources difficult to identify but including two high-voltage state-scale transmission lines compromised the ZTEM survey badly leading to unusable responses. Thus, with DOE management concurrence, the project proceeded to emphasize inversion and interpretation of the joint SubTER-FORGE MT data sets with regard to the Roosevelt Hot Springs reservoir recharge and to deep heat sources for both it and the Utah FORGE EGS project area. We also investigated the joint ZTEM-MT sampling concept with data sets from the Eleven Mile Canyon prospect area donated by the U.S. Navy (A. Sabin, PoC). Inversion of the SubTER-FORGE MT data using the HexMT 3D finite element algorithm reveals a large, low-resistivity anomaly extending sub-vertically through the depth range of the crust beneath the western Mineral Mountains. The steep conductive zone connects in the lower crust to a more tabular conductor characteristic of much of the Great Basin that generally is ascribed to current mafic magmatic underplating, hybridization and fluid release. The location of the resolved anomaly relative to the recent (0.5-0.8 Ma) eruptive centers of the Mineral Mountains implicates it as remnants of the magma body which fed these centers. This structure appears to be currently feeding heat and fluids upward into the Roosevelt Hot Springs hydrothermal system, as well as heat laterally to the FORGE project area. Separate and joint inversion models were carried out for the donated Eleven Mile Canyon MT-ZTEM data set to demonstrate concept. ZTEM only inversion showed two main alteration zones in the western portion of the project area known from geological mapping. Joint inversion including an E-W profile of MT soundings sharpened these features considerably. It also resolved in much greater detail the graben related normal faulting structure of the central project area which lies at depths exceeding the sensitivity of ZTEM alone. The sparse number of MT da-ta relative to the ZTEM required upweighting the former by a factor of several, but an exact procedure awaits future research. Our final impression is that sparse MT data can improve resolution of the subsurface over that of ZTEM alone. However, well sampled MT data are to be preferred and offer the simplicity of interpreting just one data type, and possess the superior resolution capability coming with the electric field everywhere, and from their high bandwidth.

15 GEOTHERMAL ENERGY↗

Verification Problems for Smooth Step Amplitude Load Curves in DYNA3D/Paradyn

This report documents the addition of three new verification tests in the LOADCURVE directory of the DYNA3D/Paradyn Software Quality Assurance test suite. Each test consists of a single element, where the velocities of each node are specified by either the newly added smooth step tabular load curve or another load curve option. The first test assesses the initialization and interpolation of the newly inputted load curve option through tabulated abscissa-ordinate pairs of data. The second test uses the same set of abscissa-ordinate data points and applies offset and scaling parameters available within the load curve definition. The third test defines the smooth step load curve in an original input deck, and assesses its correct redefinition using a restart file. The simulation velocities are compared to their true values at discrete points in time, and each test is verified up to numerical precision. These results confirm that the smooth step load curve option is functioning correctly and as intended.

97 MATHEMATICS AND COMPUTING↗

A Centralized AI Lakehouse Framework for Brain Tumor MRI Classification and Segmentation, University KPI Forecasting, and Water Potability Prediction

In many university and healthcare projects, models are built for very different data types such as tables, institutional time series, and medical images, but they are deployed as separate applications. In this work, that separation made testing and maintenance difficult because each module had its own pipeline and runtime requirements. This paper presents an integrated AI lakehouse-style implementation that runs three model pipelines inside one containerized backend. For medical imaging, we used MRI datasets from IEEE DataPort: a four-class classification set with 7012 images (5708 train/1304 test) and a segmentation set with 3063 image–mask pairs. The classification model (ResNet50 transfer learning) is evaluated using a proper train–validation–test protocol across multiple splits (80/10/10, 70/10/20, 60/10/30, and 10/30/60), achieving a test accuracy of 99.00% under the standard 80/10/10 split. Additionally, a patient-level evaluation is conducted using an external glioma dataset to provide a more realistic assessment without data leakage. The segmentation model (DeepLabV3-ResNet50) achieved 83.09% validation mIoU and 88.79% Dice score. For university KPI forecasting, we used annual IPEDS and NSF HERD data from 2010 to 2023 for three universities (BSU, EOU, and UAB). To examine the effect of preprocessing on forecasting performance, two case studies are conducted. In the first case, linear interpolation is applied to generate semester-level data. In the second case, the original annual data is used directly without interpolation. Random Forest regression and ARIMA models are evaluated using MAE, RMSE, MAPE, and R 2 . The results showed that interpolation improved apparent forecasting performance due to smoothing, while evaluation on the original annual data provided a more realistic assessment of model behavior. To further validate the framework on a larger dataset, an additional case study is conducted using a student dropout dataset. For water potability, we trained and compared multiple tabular classifiers on a large dataset (1,048,575 samples). A Random Forest model (100 trees, max depth 10) achieved 85.86% test accuracy and high recall for unsafe samples (0.8447). All modules are served via FastAPI and deployed together using Docker, with workflow automation routing requests to the correct endpoint. System-level benchmarking indicates that the backend maintains stable throughput and latency under concurrent requests.

97 MATHEMATICS AND COMPUTING↗

Analysis of Weather and Climate Extremes Impact on Power System Outage

This paper provides statistical analysis of the characteristics of power system outages to gain a better understanding of the impacts of the increasing severe weather conditions on the outages. 10-year historical power system outage data from the Bonneville Power Administration (BPA) were gathered together with co-located weather attributes and recorded extreme weather events in the service area, which are paired in comparable spatial and temporal scales, with a focus on each outage transmission line. Statistical frequency analysis and cross-tabular evaluation are performed to investigate the occurring frequency and duration of outages associated with extreme weather in this area of study. The study reveals that the weather-related outages can be mainly attributed to hail and thunderstorm events which correspond to up to 60% out of all failures in several transmission line types.

Ren, Huiying↗

RTN-124: Photometric Redshifts for the Vera C. Rubin Observatory Data Preview

We present the photometric redshifts (photo-z) inferred using algorithms implemented in the Redshift Assessment Infrastructure Layers (RAIL) for the NSF-DOE Vera C. Rubin Observatory Data Preview 2 (DP2). We produce a compilation of reference redshift catalog using spectroscopic, grism and many band photometric redshift dataset hosted on the LIneA Photo-z Server. We curate training and testing set for assessing the scientific and technical performance of Rubin photo-z. The algorithm applied to the object catalog are FlexZBoost, BPZ, kNN, GPz, DNF and TPz; with a combination of 6-band and 4-band photo-z depending on availability of u and y photometry. The redshift point estimates and uncertainty estimation in tabular format through the Large Survey DataBase (LSDB).

79 ASTRONOMY AND ASTROPHYSICS↗

Automatic recognition system for document digitization in nuclear power plants

With the increasing number of data-driven models in nuclear applications, large volumes of numerical data are required to accurately model and predict the health status of a plant component. However, many historical operation logs that contain useful information are not fully utilized due to the lack of a systematic approach of digitization. To overcome this issue, this study proposes an automatic pipeline for extracting information from handwritten tabular documents collected from nuclear power plants. In our pipeline, we first denoise scanned documents with morphological operations, and then extract relevant parts from individual pages using both traditional computer vision and neural network methods. Handwriting recognition is applied to obtain text and numbers. As the most challenging step is how to crop only relevant information, the main focus of our paper is to detect tables and cells from scanned handwritten documents. Here we evaluate the efficiency and accuracy of our proposed method on handwritten operational reports obtained from a real-world case study. The results demonstrate the high accuracy and practicality of our proposed method.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

OpenSesame tutorial

OpenSesame is a program for generating tabular equations of state (EOS), with capabilities for multiphase EOS construction. In this tutorial, we provide an overview of how to run OpenSesame to construct a multiphase EOS. We discuss some general features of OpenSesame, followed by a description of sample input files required for multiphase EOS construction. We also discuss how to extract data from EOS tables in order to compare to experimental data, with an example using the OpenSesame GUI. Lastly, we provide a description of how to generate ASCII-formatted EOS tables most often used by hydro code users.

97 MATHEMATICS AND COMPUTING↗

Test Information Management System (Report FY 2020 and FY 2021)

Delivery Environments supported work in Fiscal Years 2020 and 2021 to develop the Test Information Management System (TIMS). This effort involved creating a suite of tools consisting of a database and the uploading and downloading scripts. It has been demonstrated that the TIMS database is suitable for archiving raw measurements performed on various measured test articles. The TIMS database has been populated with data from the TRUST projects in FY 20 and FY 21. This included the functional data in the Sensors records as well as metadata stored in records in several other tables. The performed work demonstrated that a database archiving validation test data and integrating it with engineering analysis simulations has a great potential to increase the efficiency, responsiveness, and confidence in the modeling, simulation, and validation efforts for future (and current) weapon systems and assemblies in normal and abnormal environments. In this report, we also discuss how the records belonging to the same project should be arranged, what are the minimum requirements for linking the records using the tabular links, and we provide several recommendations to the projects engineers and the database administrator to improve the TIMS database.

97 MATHEMATICS AND COMPUTING↗

Shock compression of liquid helium to 360 GPa

Data for the shock equation of state of helium are obtained up to 360⁢G⁡Pa, nearly doubling the pressure of previous experimental measurements. The helium samples are first precompressed to 2.7 g⁡c⁢m −3 in a diamond anvil cell prior to laser-driven shock compression at the Omega Laser Facility. Time-resolved Doppler velocimetry and pyrometry reveal significant reflectivity and greater compressibility compared to the predictions of existing broad-range tabular equation of state models, which may be caused by the onset of ionization. These experimental observations, however, are largely captured with molecular-dynamics simulations based on density functional theory, affirming the ability of first-principles techniques to capture complex physics, while enabling critical insight for the behavior of warm dense helium in Jovian interiors and white dwarf atmospheres.

Physics - Plasma physics↗

Data for Rod et al., "Alternating salt and freshwater floods of coastal soils impact soil structure, hydraulic properties, and oxygen dynamics"

This dataset includes laboratory experiment data on soil structure, hydraulic properties, and oxygen dynamics associated with Rod et al. 2026 https://doi.org/10.1002/vzj2.70073. There are six data files from a lab-based flood simulation of either freshwater (FW) or alternating brackish saltwater (SW) and FW using soil cores from a coastal forest at the Smithsonian Environmental Research Center. For soil information please see the Location section of the metadata. Files include: CO2, surface chemistry, water retention, dissolved oxygen, and soil specific surface area. Each file is in CSV format and can be opened/read with any plain text tabular file reader (Microsoft Excel, R, etc.). Purpose of Experiment: To investigate how hydrologic intensification affects soil structure and oxygen dynamics, we conducted a series of laboratory-based flood simulations. After three SW-FW floods (6 floods total) there were significant changes in pore size distribution, significant redistribution of colloids, and the A-horizon became sodic. We concluded that a small number of SW flooding events can induce a measurable change in soil physical properties that directly impacts the biogeochemical dynamics.

54 ENVIRONMENTAL SCIENCES↗

INGENIOUS - Great Basin Regional Dataset Compilation

This is the regional dataset compilation for the INnovative Geothermal Exploration through Novel Investigations Of Undiscovered Systems (INGENIOUS) project. The primary goal of this project is to accelerate discoveries of new, commercially viable hidden geothermal systems while reducing the exploration and development risks for all geothermal resources. These datasets will be used in INGENIOUS as input features for predicting geothermal favorability throughout the Great Basin study area. Datasets consist of shapefiles, geotiffs, tabular spreadsheets, and metadata that describe: 2-meter temperature probe surveys, quaternary faults and volcanic features, geodetic shear and dilation models, heat flow, magnetotellurics (conductance), magnetics, gravity, paleogeothermal features (such as sinter and tufa deposits), seismicity, spring and well temperatures, spring and well aqueous geochemistry analyses, thermal conductivity, and fault slip and dilation tendency. For additional project information, see the INGENIOUS project site linked in the submission. Terms of use: These datasets are provided "as is", and the contributors assume no responsibility for any errors or omissions. The user assumes the entire risk associated with their use of these data and bears all responsibility in determining whether these data are fit for their intended use. These datasets may be redistributed with attribution (see citation information below). Please refer to the license information on this page for full licensing terms and conditions.

15 GEOTHERMAL ENERGY↗

BiG-SLiCE 2 v1.0.0

BiG-SLiCE was originally an open source Python-based command line bioinformatics software that offers a highly scalable clustering analysis on biosynthetic gene clusters (BGC) data. It allows a simultaneous analysis of millions of BGCs, exceeding the capability of other existing tools (around one hundred thousands). As a tradeoff, the clustering accuracy is relatively lower and sometimes fall short in corner cases and specific BGC classes such as the RiPPs (Ribosomally-translated, Post-translationally modified Peptides). In BiG-SLiCE V2 (developed in LBNL), the clustering algorithm has been significantly improved to deliver a much accurate result even for RiPPs and other previous corner case classes. Moreover, the speed of the overall pipeline has been improved by 50-100%. Finally, additional features were implemented to support downstream analyses of BiG-SLiCE results, such as customized tabular (TSV/CSV) and columnar (Parquet) outputs.

Kautsar, Satria↗