Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “preprocessed data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Forming a database to study reversed magnetic shear from the National Spherical Torus eXperiment using machine learning

Achieving a long-lived reversed magnetic shear (RMS) target plasma in the National Spherical Torus eXperiment Upgrade will require developing various sustainment scenarios. To help with the ongoing plasma control efforts, the development of a new analysis for the motional Stark effect (MSE) diagnostic using a machine learning algorithm, namely, MSE-ML, is described. MSE-ML will be used to identify patterns during RMS discharges, some of which suffer magnetohydrodynamic (MHD) events resulting in current redistribution and monotonic q-profiles. A database consisting of q and magnetic shear profiles is being constructed primarily based on the existing National Spherical Torus eXperiment data with equilibrium reconstructions constrained by the magnetic field pitch angle profile measured using the multi-channel MSE diagnostic. An unsupervised k-means clustering of the data is developed to study the RMS formation as a function of time. The initial clustering from the q-profiles shows significant differences in both amplitude and the duration of the RMS period. As a goal, the clustering results that detect and distinguish shots with substantial and sustained RMS are to be used as a preprocessing step in a supervised algorithm to identify the underlying conditions that lead to long-lasting improved confinement with RMS. Another aim of the MSE-ML study is to identify precursors of RMS-destroying MHD events in either derived data such as the q-profile or directly measured data such as the magnetic field pitch angle profile.

Uzun-Kaymak, I. U. (ORCID:0000000276251493)↗

Predicting the evolution of biomass bulk density through feedstock preprocessing: Discrete element modeling, regression analysis, and pilot-scale validation

Bulk density is an important material property of biomass feedstocks, influencing handling, storage, transport costs, and conversion efficiency. In this study, predictive regression models for loose and tapped bulk densities of Alamo and Cave-in-Rock switchgrass are developed using a comprehensive dataset generated via calibrated bonded-sphere discrete element method (DEM) simulations. Here, a key contribution of this study is the use of a DEM-based approach, which correlates density with moisture content and particle size distribution parameters and enables analysis across a continuous particle size range, overcoming limitations of purely experimental data. For comparison, regression models are also developed using only experimental data from pilot-scale runs at the Biomass Feedstock National User Facility at Idaho National Laboratory. Validation against pilot-scale data showed reasonable prediction accuracy for both model types, particularly for smaller particle sizes (post-secondary grinding). While the experimental model showed slightly better performance matching the validation data in some cases, the DEM-based model benefits from a much larger dataset, reduced predictor multicollinearity, and continuous parameter coverage, highlighting the utility of validated simulation models for developing robust predictive tools for biomass preprocessing applications.

09 - BIOMASS FUELS↗

Next Generation Logistics Systems for Delivering Optimal Biomass Feedstocks to Biorefining Industries in the Southeastern U.S.

The diverse portfolio of biomass sources that is available in the Southeastern U.S., including a significant supply of pine “residue”, represents a valuable strategic position for the region. Through blends formulated based on critical properties, this project will take full advantage of the range in biomass properties afforded by the portfolio to produce a consistent, high-performance feedstock for the industry, while lowering cost. Key developments being targeted to enable this potential include whole-tree transport to a state-of-the-art merchandising depot that will further access biomass from ongoing, forest industry operations. The approach will more effectively utilize the tree and distribute cost, while minimizing in-woods contamination of the woody biomass component. To implement this vision, information on the chemical composition and changes that are induced during multiple preprocessing steps (size reduction, moisture removal, densification, etc.) is needed. New NIR sensor technology will be developed for online monitoring of important biomass properties. The data will be incorporated into a statistical process control platform to improve process efficiency and meet required specifications. Advanced process models are being developed to inform the techno-economic and life-cycle assessment of the program’s impact. The new system will ultimately reduce operational risks from supply chain disruptions, and allow operation of larger-scale biorefineries.

09 BIOMASS FUELS↗

Predictive analytics of selections of russet potatoes

We explore the application of machine learning algorithms specifically to enhance the selection process of Russet potato (Solanum tuberosum L.) clones in breeding trials by predicting their suitability for advancement. This study addresses the challenge of efficiently identifying high-yield, disease-resistant, and climate-resilient potato varieties that meet processing industry standards. Leveraging manually collected data from trials in the state of Oregon, we investigate the potential of a wide variety of state-of-the-art binary classification models. The dataset includes 1086 clones, with data on 38 attributes recorded for each clone, focusing on yield, size, appearance, and frying characteristics, with several control varieties planted consistently across four Oregon regions from 2013 to 2021. We conduct a comprehensive analysis of the dataset that includes preprocessing, feature engineering, and imputation to address missing values. We focus on several key metrics such as accuracy, F1-score, and Matthews correlation coefficient (MCC) for model evaluation. The top-performing models, namely a feedforward neural network classifier (Neural Net), a histogram-based gradient boosting classifier (HGBC), and a support vector machine classifier (SVM), demonstrate consistent and significant results. To further validate our findings, we conducted a simulation study using the aims, data-generating mechanisms, estimands, methods, and performance measures (ADEMP) framework, simulating different data-generating scenarios to assess model robustness and performance through true positive, true negative, false positive, and false negative distributions, area under the receiver operating characteristic curve (AUC-ROC) and MCC. The simulation results highlight that non-linear models like SVM and HGBC consistently show higher AUC-ROC and MCC than logistic regression, thus outperforming the traditional linear model across various distributions, and emphasizing the importance of model selection and tuning in agricultural trials. Variable selection further enhances model performance and identifies influential features in predicting trial outcomes. The findings emphasize the potential of machine learning in streamlining the selection process for potato varieties, offering benefits such as increased efficiency, substantial cost savings, and judicious resource utilization. Our study contributes insights into precision agriculture and showcases the relevance of advanced technologies for informed decision-making in breeding programs.

60 APPLIED LIFE SCIENCES↗

Computer Vision Pipeline for Image Analysis for Freeze‐Fracture Electron Microscopy: Rosette Cellulose Synthase Complexes Case

In materials science, plant biology, agriculture, and environmental research, the automated analysis of high-magnification, complex microscopy images, such as those generated by freeze-fracture electron microscopy (FF-TEM), remains a critical challenge that limits the scalability of data interpretation. We present a deep learning computer vision pipeline for high-throughput detection and morphological characterization analysis of cellulose synthase complexes (CSCs, or rosettes) in FF-TEM images. The pipeline integrates preprocessing, detection, human-in-the-loop verification, and semantic segmentation to quantify features such as rosette diameter and inter-lobe spacing. The approach was trained and tested on a curated dataset of high-resolution FF-TEM micrographs of Physcomitrium patens, expanded via strategic tiling and augmentation to over 650 images. We compare YOLOv8 and YOLOv9 architectures and demonstrate that YOLOv9 achieves superior performance in both localization accuracy (mAP50-95 = 0.854) and inference speed. The resulting distributions revealed biological variability consistent with prior manual studies, validating the approach for high-throughput applications. Our results show that the pipeline achieves human-expert level accuracy while dramatically reducing analysis time, enabling scalable, reproducible structural characterization of intramembrane protein complexes. The pipeline is broadly applicable to other domains requiring precise interpretation of complex microscopy data and establishes a foundation for future artificial intelligence (AI)-assisted workflows in biological imaging.

59 BASIC BIOLOGICAL SCIENCES↗

Diffraction imaging of light induced dynamics in xenon-doped helium nanodroplets

Time-resolved wide-angle coherent diffraction imaging of individual helium nanodroplets, doped with xenon and excited with an 800 nm NIR laser pulse. Raw and preprocessed pattern, together with useful metadata, are contained in HDF5 files. Every field in the H5 container comes with a description, under the field "description", and the actual data under the field "data". For further information about the dataset see the [ArXiv article](https://arxiv.org/abs/2205.04154 ) and Bruno Langbehn's [PhD thesis](https://depositonce.tu-berlin.de/handle/11303/13189).

FERMI FEL-1↗

Diffraction imaging of light induced dynamics in xenon-doped helium nanodroplets

Time-resolved wide-angle coherent diffraction imaging of individual helium nanodroplets, doped with xenon and excited with an 800 nm NIR laser pulse. Raw and preprocessed pattern, together with useful metadata, are contained in HDF5 files. Every field in the H5 container comes with a description, under the field "description", and the actual data under the field "data". For further information about the dataset see the [ArXiv article](https://arxiv.org/abs/2205.04154 ) and Bruno Langbehn's [PhD thesis](https://depositonce.tu-berlin.de/handle/11303/13189).

FERMI FEL-1↗

Interactive Exploration of High-Dimensional Phase Diagrams

High-dimensional thermodynamic phase stability databases are becoming increasingly common due to the convergence of three recent trends: (i) the widespread interest in so-called “high-entropy” alloys, (ii) the availability of high-throughput computational assessments of phase stability in broad composition spaces and (iii) the ongoing development of ever-increasingly broad, multicomponent, multiphase CALPHAD databases. Although automated computational tools can readily process such high-dimensional data, scientists are often unable to visualize the relevant phase relations, an ability that is crucial to gaining an intuitive understanding of the stability constraints governing materials design. The present work addresses this need by providing algorithms that enable the interactive exploration of phase equilibria in high-dimensional spaces. These algorithms concentrate the complex nonlinear nonsmooth optimization needed into a preprocessing step that generates a large number of high-dimensional yet elementary graphical primitives. Furthermore, these primitives can then be cross-sectioned to yield 3-dimensional views in a computationally efficient manner that enables an interactive exploration of high-dimensional spaces. All of these operations are highly parallelizable, thus facilitating scaling of this method to large data sets.

36 MATERIALS SCIENCE↗

Super resolution for root imaging

Premise High‐resolution cameras are very helpful for plant phenotyping as their images enable tasks such as target vs. background discrimination and the measurement and analysis of fine above‐ground plant attributes. However, the acquisition of high‐resolution images of plant roots is more challenging than above‐ground data collection. An effective super‐resolution (SR) algorithm is therefore needed for overcoming the resolution limitations of sensors, reducing storage space requirements, and boosting the performance of subsequent analyses. Methods We propose an SR framework for enhancing images of plant roots using convolutional neural networks. We compare three alternatives for training the SR model: (i) training with non‐plant‐root images, (ii) training with plant‐root images, and (iii) pretraining the model with non‐plant‐root images and fine‐tuning with plant‐root images. The architectures of the SR models were based on two state‐of‐the‐art deep learning approaches: a fast SR convolutional neural network and an SR generative adversarial network. Results In our experiments, we observed that the SR models improved the quality of low‐resolution images of plant roots in an unseen data set in terms of the signal‐to‐noise ratio. We used a collection of publicly available data sets to demonstrate that the SR models outperform the basic bicubic interpolation, even when trained with non‐root data sets. Discussion The incorporation of a deep learning–based SR model in the imaging process enhances the quality of low‐resolution images of plant roots. We demonstrate that SR preprocessing boosts the performance of a machine learning system trained to separate plant roots from their background. Our segmentation experiments also show that high performance on this task can be achieved independently of the signal‐to‐noise ratio. We therefore conclude that the quality of the image enhancement depends on the desired application.

Ruiz‐Munoz, Jose F.↗

Identification of Distorted Gamma-Ray Signature Patterns Using Digital Filtering and Auto-Associative Memory Implemented with a Hopfield Neural Network

The detection and identification of radioactive sources in search applications involve analyzing passive gamma-ray emissions from high-level radioactive materials. This process uses a mobile detector-spectrometer in a complex field test environment. Recently, the use of artificial intelligence for gamma-ray spectrum analysis has shown promising results. However, challenges persist in identifying isotopic signatures from spectral measurements that may be distorted due to source shielding, random variations in natural radioactive background, or insufficient measurement time to obtain clear spectral lines. Here, this paper presents a novel intelligent signature recognition method that combines digital filtering techniques with an artificial Hopfield Neural Network (HNN). The HNN leverages auto-associative memory to store training sample patterns and match them with incoming gamma spectra from distorted sources. It restores the testing sources’ measurements by finding the closest matching signature patterns in the spectral library. Before HNN recognition, the measured spectrum undergoes preprocessing with a digital image filter to reduce fluctuations. Performance of the proposed method is evaluated using a set of gamma-ray spectra measured with a sodium iodide detector. The data collected include measurements from six pure samples: 241 Am, 60 Co, 137 Cs, 192 Ir, 239 Pu, and 235 U, which are used for training and validation (i.e. six cases). Additionally, the data set contains 24 distorted synthesized sources with various fluctuating backgrounds. Test results demonstrate the potential of the proposed method to accurately recognize the correct isotope with high precision, achieving an accuracy rate exceeding 85%. Furthermore, the proposed method exhibits superior performance compared to the conventional multiple regression fitting and simple feedforward neural network methods.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Training NuGraph2 for ICARUS

This presentation describes the process of training NuGraph2, a Graphical Neural Network for event reconstruction, on simulated ICARUS neutrino event data. This began with an investigation into filtering ICARUS spacepoint data. Then NuGraph2 was repeatedly trained on three event samples, which were used for finding optimized machine-learning parameters and to find and fix the causes of several crashes in NuGraph2 s preprocessing and training scripts.

43 PARTICLE ACCELERATORS↗

Volumetric Rendering on Wavelet-Based Adaptive Grid

Numerical modeling of physical phenomena frequently involves processes across a wide range of spatial and temporal scales. In the last two decades, the advancements in wavelet-based numerical methodologies to solve partial differential equations, combined with the unique properties of wavelet analysis to resolve localized structures of the solution on dynamically adaptive computational meshes, make it feasible to perform large-scale numerical simulations of a variety of physical systems on a dynamically adaptive computational mesh that changes both in space and time. Volumetric visualization of the solution is an essential part of scientific computing, yet the existing volumetric visualization techniques do not take full advantage of multi-resolution wavelet analysis and are not fully tailored for visualization of a compressed solution on the wavelet-based adaptive computational mesh. Our objective is to explore the alternatives for the visualization of time-dependent data on space-time varying adaptive mesh using volume rendering while capitalizing on the available sparse data representation. Two alternative formulations are explored. The first one is based on volumetric ray casting of multi-scale datasets in wavelet space. Rather than working with the wavelets at the finest possible resolution, a partial inverse wavelet transform is performed as a preprocessing step to obtain scaling functions on a uniform grid at a user-prescribed resolution. As a result, a solution in physical space is represented by a superposition of scaling functions on a coarse regular grid and wavelets on an adaptive mesh. An efficient and accurate ray casting algorithm is based just on these coarse scaling functions. Additional details are added during the ray tracing by taking an appropriate number of wavelets into account based on support overlap with the interpolation point, wavelet coefficient magnitude, and other characteristics, such as opacity accumulation (front to back ordering) and deviation from frontal viewing direction. The second approach is based on complementing of wavelet-based adaptive mesh to the traditional Adaptive Mesh Refinement (AMR) mesh. Both algorithms are illustrated and compared to the existing volume visualization software for Rayleigh-Benard thermal convection and electron density data sets in terms of rendering time and visual quality for different data compression of both wavelet-based and AMR adaptive meshes.

Vezolainen, Alexei V.↗

Adaptive continuity-preserving simplification of street networks

Street network data is widely used to study human-based activities and urban structure. Often, these data are geared towards transportation applications, which require highly granular, directed graphs that capture the complex relationships of potential traffic patterns. While this level of network detail is critical for certain fine-grained mobility models, it represents a hindrance for studies concerned with the morphology of the street network. For the latter case, street network simplification — the process of converting a highly granular input network into its most simple morphological form — is a necessary, but highly tedious preprocessing step, especially when conducted manually. In this manuscript, we develop and present a novel adaptive algorithm for simplifying street networks that is both fully automated and able to mimic results obtained through a manual simplification routine. The algorithm — available in the neatnet Python package — outperforms current state-of-the-art procedures when comparing those methods to manually, human-simplified data, while preserving network continuity.

Python↗

Real-time confinement regime detection in fusion plasmas with convolutional neural networks and high-bandwidth edge fluctuation measurements

Abstract A real-time detection of the plasma confinement regime can enable new advanced plasma control capabilities for both the access to and sustainment of enhanced confinement regimes in fusion devices. For example, a real-time indication of the confinement regime can facilitate transition to the high-performing wide-pedestal (WP) quiescent H-mode, or avoid unwanted transitions to lower confinement regimes that may induce plasma termination. To demonstrate real-time confinement regime detection, we use the 2D beam emission spectroscopy (BES) diagnostic system to capture localized density fluctuations of long wavelength turbulent modes in the edge region at a 1 MHz sampling rate. BES data from 330 discharges in either L-mode, H-mode, quiescent H (QH)-mode, or WP QH-mode were collected from the DIII-D tokamak and curated to develop a high-quality database to train a deep-learning classification model for real-time confinement detection. We utilize the 6×8 spatial configuration with a time window of 1024 µ s and recast the input to obtain spectral-like features via fast Fourier transform preprocessing. We employ a shallow 3D convolutional neural network for the multivariate time-series classification task and utilize a softmax in the final dense layer to retrieve a probability distribution over the different confinement regimes. Our model classifies the global confinement state on 44 unseen test discharges with an average F 1 score of 0.94, using only ∼1 ms snippets of BES data at a time. This activity demonstrates the feasibility for real-time data analysis of fluctuation diagnostics in future devices such as ITER, where the need for reliable and advanced plasma control is urgent.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Neural network for 3D inertial confinement fusion shell reconstruction from single radiographs

In inertial confinement fusion (ICF), x-ray radiography is a critical diagnostic for measuring implosion dynamics, which contain rich three-dimensional (3D) information. Traditional methods for reconstructing 3D volumes from 2D radiographs, such as filtered backprojection, require radiographs from at least two different angles or lines of sight (LOS). In ICF experiments, the space for diagnostics is limited, and cameras that can operate on fast timescales are expensive to implement, limiting the number of projections that can be acquired. To improve the imaging quality as a result of this limitation, convolutional neural networks (CNNs) have recently been shown to be capable of producing 3D models from visible light images or medical x-ray images rendered by volumetric computed tomography. Here, we propose a CNN to reconstruct 3D ICF spherical shells from single radiographs. We also examine the sensitivity of the 3D reconstruction to different illumination models using preprocessing techniques such as pseudo-flatfielding. To resolve the issue of the lack of 3D supervision, we show that training the CNN utilizing synthetic radiographs produced by known simulation methods allows for reconstruction of experimental data as long as the experimental data are similar to the synthetic data. We also show that the CNN allows for 3D reconstruction of shells that possess low mode asymmetries. Further comparisons of the 3D reconstructions with direct multiple LOS measurements are justified.

3D Reconstruction↗

Kβ X-ray Emission Spectroscopy as a Probe of Cu(I) Sites: Application to the Cu(I) Site in Preprocessed Galactose Oxidase

Cu(I) active sites in metalloproteins are involved in O 2 activation, but their O 2 reactivity is difficult to study due to the Cu(I) d 10 closed shell which precludes the use of conventional spectroscopic methods. Kβ X-ray emission spectroscopy (XES) is a promising technique for investigating Cu(I) sites as it detects photons emitted by electronic transitions from occupied orbitals. Here, we demonstrate the utility of Kβ XES in probing Cu(I) sites in model complexes and a metalloprotein. Using Cu(I)Cl, emission features from double-ionization (DI) states are identified using varying incident X-ray photon energies, and a reasonable method to correct the data to remove DI contributions is presented. Kβ XES spectra of Cu(I) model complexes, having biologically relevant N/S ligands and different coordination numbers, are compared and analyzed, with the aid of density functional theory (DFT) calculations, to evaluate the sensitivity of the spectral features to the ligand environment. While the low-energy Kβ 2,5 emission feature reflects the ionization energy of ligand np valence orbitals, the high-energy Kβ 2,5 emission feature corresponds to transitions from molecular orbitals (MOs) having mainly Cu 3d character with the intensities determined by ligand-mediated d–p mixing. A Kβ XES spectrum of the Cu(I) site in preprocessed galactose oxidase (GO pre ) supports the 1Tyr/2His structural model that was determined by our previous X-ray absorption spectroscopy and DFT study. Finally, the high-energy Kβ 2,5 emission feature in the Cu(I)-GO pre data has information about the MO containing mostly Cu 3d x 2 –y 2 character that is the frontier molecular orbital (FMO) for O 2 activation, which shows the potential of Kβ XES in probing the Cu(I) FMO associated with small-molecule activation in metalloproteins.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

NLR HPC Kestrel Jobs Data

Overview: Anonymized job-level records from the Kestrel HPC system at the National Laboratory of the Rockies (NLR). Each record represents a Slurm batch job with scheduling metadata, resource requests, utilization, energy estimates, and efficiency metrics. Sensitive fields (user, account, job name, submit line, working directory, submit script, and job type) are replaced with 7-character cryptographic hashes. System & Timeframe: Kestrel is located at the NLR campus. Standard compute nodes have 104 cores and 256 GB RAM; bigmem nodes have 2,000 GB. GPU nodes (gpu-h100 partition) use NVIDIA H100 GPUs. Data covers jobs submitted August 2023 through December 2025. Funding provided by the U.S. Department of Energy, EERE. Files: esif.hpc.kestrel.job-anon.zip — Anonymized job records (Hive-partitioned Parquet) datacard.md — Full dataset documentation ~11 million rows, 50 variables. Readable with PyArrow, pandas, DuckDB, Apache Spark, or any Parquet-compatible tool. Data Collection: Jobs collected via sacct with timezone-aware export (SLURM_TIME_FORMAT="%Y-%m-%dT%H:%M:%S%z"), loaded into PostgreSQL. Calculated columns updated via database triggers and batch functions. All timestamps use timestamptz and correctly handle DST transitions. Preprocessing: Anonymization of name, user, account, submit_line, work_dir, submit_script, and job_type via 7-char hex hashes Derived columns: queue_wait, cpu_eff, max/min/avg_mem_eff, energy estimates Simplified job state mapping (e.g., "CANCELLED by 132357" → "CANCELLED") Boolean flags: python_job, reframe_job Temporal decomposition: year, month, day, day_of_week, hour, minute from submit_time Shared node tracking: shared_job_count, nodes_shared, jobs_shared Key Variables: Scheduling: job_id, partition, state_simple, submit_time, start_time, end_time, queue_wait Resources: nodes_req/used, processors_req/used, memory_req, wallclock_req/used, gpus_requested Efficiency: cpu_eff, max/min/avg_mem_eff Energy: cpu_energy_tdp_estimated_max/used_watt_hours, consumed_energy_raw_joules, consumed_energy_raw_watt_hours Sharing: shared_job_count, nodes_shared, jobs_shared Partitions: short, standard, debug, gpu-h100 Job States: CANCELLED, COMPLETED, FAILED, PENDING, RUNNING QoS Levels: normal, high Important Notes: Timestamps include timezone offsets; DST transitions are handled correctly, though adding intervals across DST boundaries requires offset adjustment shared_job_count reflects physical node co-residency, not use of the shared partition Job step records and raw Slurm JSONB fields are excluded Do not attempt to re-identify individuals from hashed fields

97 MATHEMATICS AND COMPUTING↗

NLR HPC Eagle Jobs Data and Additional Energy Metrics

Overview: Anonymized job-level records from the Eagle high-performance computing (HPC) system at the National Laboratory of the Rockies (NLR). Each record represents a Slurm batch job with scheduling metadata, resource requests, resource utilization, CPU/GPU energy consumption, and efficiency metrics. Sensitive fields (user, account, job name) are replaced with cryptographic hashes. System & Timeframe: Eagle was a 2,000-node, 8-petaflop system operated at NLR from 2019–2024. Data covers the full operational lifetime of the system. Slurm data was processed nightly; timestamps are in Mountain Time. Funding provided by the U.S. Department of Energy, EERE. Files: esif.hpc.eagle.job-anon.zip — Core anonymized job records (Hive-partitioned Parquet) esif.hpc.eagle.job-anon-energy-metrics.zip — Same records with additional iLO and Ganglia energy metrics datacard.md — Full dataset documentation ~13.8 million rows, 62 variables. Readable with PyArrow, pandas, DuckDB, Apache Spark, or any Parquet-compatible tool. Data Collection: Jobs collected via sacct through a pipeline: Eagle Jobs API → Redpanda → StreamSets → HPCMON API → PostgreSQL. Node-level power from iLO (HP Integrated Lights-Out); GPU power from Ganglia monitoring, joined to jobs via node lists and time ranges. Preprocessing: Anonymization of name, user, and account fields via cryptographic hashing Derived columns: queue_wait, cpu_eff, max_mem_eff Simplified job state mapping (e.g., "CANCELLED BY 12345" → "CANCELLED") QoS accounting rules (buy-in, standby, or Slurm QoS value) CPU energy estimated from TDP (200W, Intel Xeon Gold 6154, 18 cores) Timezone-aware columns (_tz) sourced from LEX accounting database to correctly handle DST transitions Key Variables: Scheduling: job_id, partition, state_simple, submit_time_tz, start_time_tz, end_time_tz, queue_waitResources: nodes_req/used, processors_req/used, memory_req, wallclock_req/used, gpus_requested Efficiency: cpu_eff, max_mem_eff Energy: cpu_energy_tdp_estimated_max/used_watt_hours, node_energy_total_watt_hours (iLO), gpu0/1_energy_total_watt_hours (Ganglia) Partitions: bigmem, bigmem-8600, bigscratch, csc, dav, ddn, debug, gpu, haswell, long, mono, short, standard Job States: CANCELLED, COMPLETED, FAILED, NODE_FAIL, OUT_OF_MEMORY, PENDING, RUNNING, TIMEOUT QoS Levels: Unknown, normal, buy-in, debug, penalty, high, standby Important Notes: Non-_tz timestamp columns may be off by one hour across DST boundaries; use _tz columns for time difference calculations Energy fields are null for jobs without monitoring coverage Job step records and raw Slurm JSONB fields are excluded from this extract Do not attempt to re-identify individuals from hashed fields

97 MATHEMATICS AND COMPUTING↗