Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “preprocessing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Genome collection processing for “Conserved upper thermal limits and small safety margins in soil copiotrophic bacteria”

We extracted the genomic DNA of 400 randomly selected isolates using a Quick-DNA Microprep Kit (Zymo Research D3020) according to the manufacturer’s protocol. We then submitted the extracted gDNA samples for short-read Illumina sequencing (200 Mbp) at SeqCoast Genomics (Portsmouth, NH, USA). After preprocessing the sequences using Trimmommatic (Bolger et al. 2014), we assembled the genomes using SPADES (Bankevich et al. 2012) and checked the quality of each assembly using QUAST (Gurevich et al. 2013). We processed the genome assemblies using a KBase (v1.4.0) pipeline (Allen et al. 2017; Arkin et al. 2018). Briefly, we used DRAM (v0.1.2) with default settings to annotate the genome assemblies. We then evaluated genome quality and possible contamination levels using CheckM (v1.0.18) (Parks et al. 2015) and retained genomes with completeness above 98% and contamination below 5% (n = 354), following the authors' guidelines. We then obtained taxonomic assignments for all remaining isolates using the Genome Taxonomy Database tool GTDB-Tk (v2.3.2, database version r214) (Chaumeil et al. 2019). We constructed a phylogenetic tree using the tool SpeciesTree (v2.2.0). We then trimmed the tree (using Trim SpeciesTree to GenomeSet- v1.4.0), retaining only tips within our collection with measured thermal performance.

59 BASIC BIOLOGICAL SCIENCES↗

Addressing GPU memory limitations for Graph Neural Networks in High-Energy Physics applications

Introduction Reconstructing low-level particle tracks in neutrino physics can address some of the most fundamental questions about the universe. However, processing petabytes of raw data using deep learning techniques poses a challenging problem in the field of High Energy Physics (HEP). In the Exa.TrkX Project, an illustrative HEP application, preprocessed simulation data is fed into a state-of-art Graph Neural Network (GNN) model, accelerated by GPUs. However, limited GPU memory often leads to Out-of-Memory (OOM) exceptions during training, due to the large size of models and datasets. This problem is exacerbated when deploying models on High-Performance Computing (HPC) systems designed for large-scale applications. Methods We observe a high workload imbalance issue during GNN model training caused by the irregular sizes of input graph samples in HEP datasets, contributing to OOM exceptions. We aim to scale GNNs on HPC systems, by prioritizing workload balance in graph inputs while maintaining model accuracy. Our paper introduces diverse balancing strategies aimed at decreasing the maximum GPU memory footprint and avoiding the OOM exception, across various datasets. Results Our experiments showcase memory reduction of up to 32.14% compared to the baseline. We also demonstrate the proposed strategies can avoid OOM in application. Additionally, we create a distributed multi-GPU implementation using these samplers to demonstrate the scalability of these techniques on the HEP dataset. Discussion By assessing the performance of these strategies as data loading samplers across multiple datasets, we can gauge their effectiveness in both single-GPU and distributed environments. Our experiments, conducted on datasets of varying sizes and across multiple GPUs, broaden the applicability of our work to various GNN applications that handle input datasets with irregular graph sizes.

Lee, Claire Songhyun↗

Applications of LIF to Document Natural Variability of Chlorophyll Content and Cu Uptake in Moss

Chlorophyll has long been used as a natural indicator of plant health and photosynthetic efficiency. Laser-induced fluorescence (LIF) is an emerging technique for understanding broad spectrum organic processes and has more recently been used to monitor chlorophyll response in plants. Previous work has focused on developing a LIF technique for imaging moss mats to identify metal contamination with the current focus shifting toward application to moss fronds and aiding sample collection for chemical analysis. Two laser systems (CoCoBi a Nd:YGa pulsed laser system and Chl-SL with two blue continuous semiconductor diodes) were used to collect images of moss fronds exposed to increasing levels of Cu (1, 10, and 100 nmol/cm 2 ) using a CMOS camera. The best methods for the preprocessing of images were conducted before the analysis of fluorescence signatures were compared to a control. The Chl-SL system performed better than the CoCoBi, with dynamic time warping (DTW) proving the most effective for image analysis. Manual thresholding to remove lower decimal code values improved the data distributions and proved whether using one or two fronds in an image was more advantageous. A higher DTW difference from the control correlated to lower chlorophyll a/b ratios and a higher metal content, indicating that LIF, with the aid of image processing, can be an effective technique for identifying Cu contamination shortly after an event.

59 BASIC BIOLOGICAL SCIENCES↗

A Centralized AI Lakehouse Framework for Brain Tumor MRI Classification and Segmentation, University KPI Forecasting, and Water Potability Prediction

In many university and healthcare projects, models are built for very different data types such as tables, institutional time series, and medical images, but they are deployed as separate applications. In this work, that separation made testing and maintenance difficult because each module had its own pipeline and runtime requirements. This paper presents an integrated AI lakehouse-style implementation that runs three model pipelines inside one containerized backend. For medical imaging, we used MRI datasets from IEEE DataPort: a four-class classification set with 7012 images (5708 train/1304 test) and a segmentation set with 3063 image–mask pairs. The classification model (ResNet50 transfer learning) is evaluated using a proper train–validation–test protocol across multiple splits (80/10/10, 70/10/20, 60/10/30, and 10/30/60), achieving a test accuracy of 99.00% under the standard 80/10/10 split. Additionally, a patient-level evaluation is conducted using an external glioma dataset to provide a more realistic assessment without data leakage. The segmentation model (DeepLabV3-ResNet50) achieved 83.09% validation mIoU and 88.79% Dice score. For university KPI forecasting, we used annual IPEDS and NSF HERD data from 2010 to 2023 for three universities (BSU, EOU, and UAB). To examine the effect of preprocessing on forecasting performance, two case studies are conducted. In the first case, linear interpolation is applied to generate semester-level data. In the second case, the original annual data is used directly without interpolation. Random Forest regression and ARIMA models are evaluated using MAE, RMSE, MAPE, and R 2 . The results showed that interpolation improved apparent forecasting performance due to smoothing, while evaluation on the original annual data provided a more realistic assessment of model behavior. To further validate the framework on a larger dataset, an additional case study is conducted using a student dropout dataset. For water potability, we trained and compared multiple tabular classifiers on a large dataset (1,048,575 samples). A Random Forest model (100 trees, max depth 10) achieved 85.86% test accuracy and high recall for unsafe samples (0.8447). All modules are served via FastAPI and deployed together using Docker, with workflow automation routing requests to the correct endpoint. System-level benchmarking indicates that the backend maintains stable throughput and latency under concurrent requests.

97 MATHEMATICS AND COMPUTING↗

H I Depletion Begins Well Beyond the Virial Radius: A FAST Stacking Study of 36 Galaxy Clusters to 5 × R 200

Abstract We present a stacking study of the neutral atomic hydrogen (H i ) content in and around 36 local galaxy clusters at z < 0.07, using a combination of the FAST All Sky H i survey (FASHI) and the extensive spectroscopic catalog mainly from the Dark Energy Spectroscopic Instrument (DESI). We employ spectral stacking techniques to probe the average H i mass and HI-to-stellar mass ratio ( M HI / M * ) for member galaxies down to stellar masses of M * ∼ 10 9 M ⊙ , spanning a projected cluster-centric distance of up to 5 R 200 . Our analysis reveals a pronounced environmental effect; both M HI and M HI / M * decrease steadily toward the cluster center, dropping by ∼0.5 dex on average from the outskirts to the core. Crucially, we find that M HI / M * of galaxies remain lower than the field galaxies even at the 5 R 200 . This provides direct, statistical evidence for substantial gas stripping and preprocessing in the cluster outskirts, likely occurring in infalling groups and large-scale filaments. By further splitting the sample by g − r color, we show that the H i deficiency persists at fixed galaxy color; even the bluest cluster members exhibit ∼0.5 dex lower M HI / M * than field galaxies of similar color, reflecting environmental effects on the cold gas reservoir prior to full optical transformation. The total H i mass within clusters and their outskirts agrees broadly with predictions from cosmological simulations. Our results underscore the critical role of the extended cluster environment in quenching galaxies by depleting their cold gas reservoirs well before they enter the dense cluster core.

Cheng, Cheng [Chinese Academy of Sciences South Am↗

NLR HPC Kestrel Jobs Data

Overview: Anonymized job-level records from the Kestrel HPC system at the National Laboratory of the Rockies (NLR). Each record represents a Slurm batch job with scheduling metadata, resource requests, utilization, energy estimates, and efficiency metrics. Sensitive fields (user, account, job name, submit line, working directory, submit script, and job type) are replaced with 7-character cryptographic hashes. System & Timeframe: Kestrel is located at the NLR campus. Standard compute nodes have 104 cores and 256 GB RAM; bigmem nodes have 2,000 GB. GPU nodes (gpu-h100 partition) use NVIDIA H100 GPUs. Data covers jobs submitted August 2023 through December 2025. Funding provided by the U.S. Department of Energy, EERE. Files: esif.hpc.kestrel.job-anon.zip — Anonymized job records (Hive-partitioned Parquet) datacard.md — Full dataset documentation ~11 million rows, 50 variables. Readable with PyArrow, pandas, DuckDB, Apache Spark, or any Parquet-compatible tool. Data Collection: Jobs collected via sacct with timezone-aware export (SLURM_TIME_FORMAT="%Y-%m-%dT%H:%M:%S%z"), loaded into PostgreSQL. Calculated columns updated via database triggers and batch functions. All timestamps use timestamptz and correctly handle DST transitions. Preprocessing: Anonymization of name, user, account, submit_line, work_dir, submit_script, and job_type via 7-char hex hashes Derived columns: queue_wait, cpu_eff, max/min/avg_mem_eff, energy estimates Simplified job state mapping (e.g., "CANCELLED by 132357" → "CANCELLED") Boolean flags: python_job, reframe_job Temporal decomposition: year, month, day, day_of_week, hour, minute from submit_time Shared node tracking: shared_job_count, nodes_shared, jobs_shared Key Variables: Scheduling: job_id, partition, state_simple, submit_time, start_time, end_time, queue_wait Resources: nodes_req/used, processors_req/used, memory_req, wallclock_req/used, gpus_requested Efficiency: cpu_eff, max/min/avg_mem_eff Energy: cpu_energy_tdp_estimated_max/used_watt_hours, consumed_energy_raw_joules, consumed_energy_raw_watt_hours Sharing: shared_job_count, nodes_shared, jobs_shared Partitions: short, standard, debug, gpu-h100 Job States: CANCELLED, COMPLETED, FAILED, PENDING, RUNNING QoS Levels: normal, high Important Notes: Timestamps include timezone offsets; DST transitions are handled correctly, though adding intervals across DST boundaries requires offset adjustment shared_job_count reflects physical node co-residency, not use of the shared partition Job step records and raw Slurm JSONB fields are excluded Do not attempt to re-identify individuals from hashed fields

97 MATHEMATICS AND COMPUTING↗

NLR HPC Eagle Jobs Data and Additional Energy Metrics

Overview: Anonymized job-level records from the Eagle high-performance computing (HPC) system at the National Laboratory of the Rockies (NLR). Each record represents a Slurm batch job with scheduling metadata, resource requests, resource utilization, CPU/GPU energy consumption, and efficiency metrics. Sensitive fields (user, account, job name) are replaced with cryptographic hashes. System & Timeframe: Eagle was a 2,000-node, 8-petaflop system operated at NLR from 2019–2024. Data covers the full operational lifetime of the system. Slurm data was processed nightly; timestamps are in Mountain Time. Funding provided by the U.S. Department of Energy, EERE. Files: esif.hpc.eagle.job-anon.zip — Core anonymized job records (Hive-partitioned Parquet) esif.hpc.eagle.job-anon-energy-metrics.zip — Same records with additional iLO and Ganglia energy metrics datacard.md — Full dataset documentation ~13.8 million rows, 62 variables. Readable with PyArrow, pandas, DuckDB, Apache Spark, or any Parquet-compatible tool. Data Collection: Jobs collected via sacct through a pipeline: Eagle Jobs API → Redpanda → StreamSets → HPCMON API → PostgreSQL. Node-level power from iLO (HP Integrated Lights-Out); GPU power from Ganglia monitoring, joined to jobs via node lists and time ranges. Preprocessing: Anonymization of name, user, and account fields via cryptographic hashing Derived columns: queue_wait, cpu_eff, max_mem_eff Simplified job state mapping (e.g., "CANCELLED BY 12345" → "CANCELLED") QoS accounting rules (buy-in, standby, or Slurm QoS value) CPU energy estimated from TDP (200W, Intel Xeon Gold 6154, 18 cores) Timezone-aware columns (_tz) sourced from LEX accounting database to correctly handle DST transitions Key Variables: Scheduling: job_id, partition, state_simple, submit_time_tz, start_time_tz, end_time_tz, queue_waitResources: nodes_req/used, processors_req/used, memory_req, wallclock_req/used, gpus_requested Efficiency: cpu_eff, max_mem_eff Energy: cpu_energy_tdp_estimated_max/used_watt_hours, node_energy_total_watt_hours (iLO), gpu0/1_energy_total_watt_hours (Ganglia) Partitions: bigmem, bigmem-8600, bigscratch, csc, dav, ddn, debug, gpu, haswell, long, mono, short, standard Job States: CANCELLED, COMPLETED, FAILED, NODE_FAIL, OUT_OF_MEMORY, PENDING, RUNNING, TIMEOUT QoS Levels: Unknown, normal, buy-in, debug, penalty, high, standby Important Notes: Non-_tz timestamp columns may be off by one hour across DST boundaries; use _tz columns for time difference calculations Energy fields are null for jobs without monitoring coverage Job step records and raw Slurm JSONB fields are excluded from this extract Do not attempt to re-identify individuals from hashed fields

97 MATHEMATICS AND COMPUTING↗

Herbaceous Feedstock 2022 State of Technology Report

The U.S. Department of Energy promotes production of advanced liquid transportation fuels from lignocellulosic biomass by funding fundamental and applied research that advances the state of technology (SOT). As part of its involvement in this mission, Idaho National Laboratory completes an annual SOT report for nth-plant and 1st-plant herbaceous biomass feedstock logistics. The purpose of the SOT is to provide the status of feedstock supply system technology development for herbaceous biomass to biofuels relative to technical targets and cost goals from specific design cases, based on data and experimental results. Although conventional feedstock supply systems form the backbone of the emerging biofuels industry, they have limitations that restrict widespread implementation on a national scale. To meet the demands of the future industry, the feedstock supply system must shift from the conventional system to what has been termed “advanced” supply systems. In advanced designs, a distributed network of aggregation and processing centers, termed “depots,” are employed near the points of biomass production (i.e., the field or forest) to reduce feedstock variability and produce feedstocks of a uniform format, moving toward biomass commoditization. The 2022 Herbaceous SOT is part of a vision of achieving an implemented advanced feedstock supply system, which produces a stable, tradable commodity at the decentralized distributed depot. It utilizes feedstock fractionation by incorporating technologies that can separate the biomass into its anatomical fractions (leaves, husks, stems and cobs) to reduce impurities and produce fractions that satisfy downstream quality considerations. By using a series of air classification steps, this strategy can reduce the extrinsic ash in corn stover and produce enriched tissue fractions that can be blended to a conversion specification or converted individually in optimized biochemical conversion campaigns. Additionally, a majority of the leaves (which do not meet the quality specification) are separated out early and can be supplied to alternate markets. The 2022 Herbaceous SOT incorporates an advanced biomass fractionation and processing system to produce pellets enriched tissues from three-pass corn stover. The resulting enriched pellets are delivered to the biorefinery individually where they can be blended to a specification or converted in campaigns where the conditions are optimized for each tissue. Unused fractions can be sent to a a midstream market or to a different conversion process that is better suited to their properties to offset the cost of the delivered feedstock. The main benefits from the proposed system can be summarized as: (1) $6.86/dry ton (2016$) lower cost for the air classification due to elimination of the requirement to discard the high ash lights fraction; (2) $1.56/dry ton lower delivered cost by selling the unsuitable leaf fraction into the feed market as a midstream co-product (assuming a selling price that is 11% higher than their cost of production); (3) 0.98% increase in carbohydrate content (from 60.16% to 61.14%); and (4) 0.97% decrease in ash content (from 6.00% to 5.03%) compared to the 2021 Herbaceous SOT. Overall, the 2022 nth-plant Herbaceous SOT predicts a modeled delivered feedstock cost of $78.64/dry ton (2016$) if it is assumed that the enriched leaf fraction is sold at its production cost; this is a slight increase of $0.43/dry ton increase from the 2021 Herbaceous SOT nth-Supply case cost. The increased cost derived from a $0.38/dry ton increase in transportation and handling cost to procure more biomass (to replace the enriched leaf fraction that was not delivered to the biorefinery. The total preprocessing cost was $0.27/dry ton higher than the 2021 result because of updates to energy consumption, purchasing price and dry matter loss data for the rotary shear ($3.00/dry ton increase) and the pelleting mill ($4.52/dry ton increase). The data utilized were generated in pilot-scale tests in the Biomass Feedstock National User Facility (BFNUF) at INL and at Forest Concepts, including tests for rotary shear and pelleting of the air classified fractions. A greenhouse gas emissions analysis was performed by Argonne National Laboratory using the most up to date version of the Greenhouse Gases, Regulated Emissions, and Energy use in Transportation model (GREET®). The analysis showed an increase of 17.34 kg CO2e/dry ton from the 2021 SOT (67.71 kg CO2e/ton in the 2021 Herbaceous SOT to 85.05 kg CO2e/ton in the 2022 Herbaceous SOT). The net increase is primarily attributed to increased energy consumption in pelleting mill.

09 BIOMASS FUELS↗

Investigating the Effect of Water on the Mechanical Properties of Cellulose from Multiscale Molecular Dynamics Simulations

Classical molecular dynamics (MD) simulations provide insight into the structure and physicochemical properties of materials with atomic resolution. However, the length and time scales accessible to atomistic MD are orders of magnitude smaller than many relevant processes such as the response of a bulk material to experimentally accessible strain rates, which presents challenges when comparing models to experimental measurements. Bottom-up coarse-graining provides a means for systematically mapping atomistic information to lower resolution models to increase the length and time scales achievable by simulation. Cellulose is an abundant carbohydrate biopolymer with applications to many fields of research, such as materials science and renewable energy, due to its desirable mechanical properties and viability for conversion into biofuel. The effect of moisture content on the Young's modulus of cellulose is of special interest due to its native environment often being in the hydrated secondary plant cell wall and the grinding energy requirements for biomass feedstock preprocessing. The current work investigates the effects of water solvent on the Young's modulus of cellulose calculated from coarse-grained MD mechanical stress simulations. The coarse-grained model was parametrized from atomistic MD calculations of cellulose-cellulose potentials of mean force using umbrella sampling techniques under vacuum and solvated conditions. The Young's moduli of the coarse-grained cellulose assemblies parametrized from cellulose in vacuum or solvated in water were computed via mechanical stress simulations to highlight the importance of capturing solvent interactions for modeling the mechanical behavior of cellulose.

BASIC BIOLOGICAL SCIENCES,RADIATION PROTECTION AND↗

PhotonIDs: ML-Powered Photon Identification System for Dark Count Elimination

Reliable single photon detection is the foundation for practical quantum communication and networking. However, today's superconducting nanowire single photon detector(SNSPD) inherently fails to distinguish between genuine photon events and dark counts, leading to degraded fidelity in long-distance quantum communication. In this work, we introduce PhotonIDs, a machine learning-powered photon identification system that is the first end-to-end solution for real-time discrimination between photons and dark count based on full SNSPD readout signal waveform analysis. PhotonIDs ~demonstrates: 1) an FPGA-based high-speed data acquisition platform that selectively captures the full waveform of signal only while filtering out the background data in real time; 2) an efficient signal preprocessing pipeline, and a novel pseudo-position metric that is derived from the physical temporal-spatial features of each detected event; 3) a hybrid machine learning model with near 98% accuracy achieved on photon/dark count classification. Additionally, proposed PhotonIDs ~ is evaluated on the dark count elimination performance with two real-world case studies: (1) 20 km quantum link, and (2) Erbium ion-based photon emission system. Our result demonstrates that PhotonIDs ~could improve more than 31.2 times of signal-noise-ratio~(SNR) on dark count elimination. PhotonIDs ~ marks a step forward in noise-resilient quantum communication infrastructure.

Linne, Karl C. [Chicago U.] (ORCID:000900091870358↗

An Evaluation of Representation Learning Methods in Particle Physics Foundation Models

We present a systematic evaluation of representation learning objectives for particle physics within a unified framework. Our study employs a shared transformer-based particle-cloud encoder with standardized preprocessing, matched sampling, and a consistent evaluation protocol on a jet classification dataset. We compare contrastive (supervised and self-supervised), masked particle modeling, and generative reconstruction objectives under a common training regimen. In addition, we introduce targeted supervised architectural modifications that achieve state-of-the-art performance on benchmark evaluations. This controlled comparison isolates the contributions of the learning objective, highlights their respective strengths and limitations, and provides reproducible baselines. We position this work as a reference point for the future development of foundation models in particle physics, enabling more transparent and robust progress across the community.

Chen, Michael [Caltech]↗

LUNA: LUT-Based Neural Architecture for Fast and Low-Cost Qubit Readout

Qubit readout is a critical operation in quantum computing systems, which maps the analog response of qubits into discrete classical states. Deep neural networks (DNNs) have recently emerged as a promising solution to improve readout accuracy . Prior hardware implementations of DNN-based readout are resource-intensive and suffer from high inference latency, limiting their practical use in low-latency decoding and quantum error correction (QEC) loops. This paper proposes LUNA, a fast and efficient superconducting qubit readout accelerator that combines low-cost integrator-based preprocessing with Look-Up Table (LUT) based neural networks for classification. The architecture uses simple integrators for dimensionality reduction with minimal hardware overhead, and employs LogicNets (DNNs synthesized into LUT logic) to drastically reduce resource usage while enabling ultra-low-latency inference. We integrate this with a differential evolution based exploration and optimization framework to identify high-quality design points. Our results show up to a 10.95x reduction in area and 30% lower latency with little to no loss in fidelity compared to the state-of-the-art. LUNA enables scalable, low-footprint, and high-speed qubit readout, supporting the development of larger and more reliable quantum computing systems.

Farooq, M. A. [Arizona State U., Tempe]↗

A Knowledge Graph Approach to Analyze Systems and Assets Health

Nuclear power plants collect large amounts of equipment reliability data elements that contain information on the statuses of component, assets, and systems. All these data elements precisely record asset and system performance and health throughout the lifecycle of those assets and systems. However, several challenges have proved to be roadblocks to this process. While some of these challenges are technical in nature (i.e., data are often distributed over several physical servers or databases), others are conceptual in nature (i.e., data elements come in different formats, numeric or textual), and measured values have different scales (e.g., vibration spectra and oil temperature). This paper directly focuses on the integration of numeric and textual data elements in order to assist plant system engineers in analyzing equipment reliability data. This task begins with preprocessing the data by extracting knowledge from textual data via natural language processing methods and quantifying system, asset, and component health based on numeric data. We then employed model-based system engineering (MBSE) models of systems and assets to identify their architecture and functional (i.e., cause and effect) relations. Data elements were then associated with a single MBSE graph element, based on their nature. This bonding of MBSE models and data elements constitutes a first-of-its-kind knowledge graph of a nuclear power plants system, with data elements being organized in a structured manner that enables system engineers to identify cause-effect trends in data elements and carry out appropriate actions in response.

97 - MATHEMATICS AND COMPUTING↗

Air Classification of Forestry Residues for Fast Pyrolysis

Understanding critical biomass attributes through efficient fractionation is crucial for advancing sustainable pyrolysis for renewable energy and chemical production. This study investigates the intricate relationship between biomass preprocessing and pyrolysis product yields, employing the air classification technique for the treatment of loblolly pine residues with varying moisture content. A comprehensive exploration of the physicochemical properties of air-classified loblolly pine informs a sophisticated pyrolysis simulation model. Given the complex and multifaceted nature of biomass pyrolysis, operating across diverse temporal and spatial scales, a pyrolysis kinetics-based CFD–DEM simulation method is employed to predict product yields. Results showed that the elevated moisture content amplifies particle adhesiveness, necessitating augmented air velocities for effective separation, thereby influencing the efficiency of the separation process. While carbon and hydrogen contents exhibit relative stability across diverse moisture contents and blower frequencies, the oxygen content undergoes noticeable changes. For example, the oxygen contents were measured as 29.2 and 38.6 wt% in the light fraction of 30% moisture content sample at blower frequencies of 10 and 20 Hz, respectively. An intriguing finding emerges from pyrolysis simulation, indicating that a lower blower frequency in air classification moderately enhances bio-oil yield and significantly improves its quality, particularly in terms of water content. For instance, the water content in the bio-oil was about 1.5% and 10% in the heavy and light fractions, respectively from 10% moisture sample under 15 Hz blower frequency.

09 - BIOMASS FUELS↗

Design Choices in Anomaly Detection for Industrial Control Systems: Insights from Gas Pipeline Data

Industrial control systems (ICS) remain vulnerable to increasingly sophisticated cyberattacks, yet evaluating anomaly detection models in these environments is challenging due to temporal dependencies, missing-not-at-random patterns, and extremely imbalanced datasets. These factors make common practices—especially random data splits and naïve imputation—prone to severe temporal leakage, which can inflate reported performance and obscure real-world limitations. In this work, we systematically examine classical machine learning models, temporal deep learning architecture, and tensor-decomposition–based methods on a gas-pipeline dataset using a fully temporally separated evaluation pipeline designed to mimic realistic deployment conditions. Our findings show that proper temporal handling and MNAR-aware preprocessing significantly alter the relative performance of popular anomaly-detection methods, providing practical guidance for designing reliable, leakage-resistant ICS intrusion-detection systems.

97 MATHEMATICS AND COMPUTING↗

Demystifying Piecewise and Localized Scatter Correction Methods

Multiplicative scatter is a common source of noise in near-infrared spectroscopy and other related instrumental techniques. A wide variety of methods are commonly used for the correction of multiplicative scatter. However, the majority of such methods assume that the parameters that describe the scatter are constant throughout the measured spectrum, which is often not the case. This work investigates a family of methods that perform scatter correction using local regions of neighboring wavelengths in order to better account for wavelength-dependent scattering. The methods in question are piecewise standard normal variate, localized standard normal variate, piecewise multiplicative scatter correction, and localized multiplicative scatter correction. This work describes the theoretical and algorithmic foundations of the family of local region-based scatter correction methods and compares their application and optimization at a qualitative and quantitative level using several datasets.

NIR spectroscopy↗

Enhancing the flowability of woody biomass slurries in wet biorefineries

Feeding wet lignocellulosic biomass (e.g., softwood and hardwood) slurries into high-pressure, high-temperature reactors at an industrially relevant scale presents significant challenges, such as process equipment plugging. Here, in this study, we investigate the possibility of improving the flowability of biomass slurries in an industrial scale wet biorefinery by tuning the physical and chemical characteristics of the biomass particles. To understand the effects of the chemical characteristics of woody biomass on flowability, cellulose pulp particles are produced from pulp sheets via knife milling, pelletizing, and crumbling. These cellulose pulp particles are processed for acid hydrolysis dehydration (AHDH) at one –tonne-per day pilot scale. The levulinic acid yield (a product of cellulose AHDH) and the flowability of the biomass particles are compared using market pulp, 1 mm crumbled pine softwood, a blend of 2 mm hammer-milled pine softwood with 10 wt.% bark, and 2 mm hammer-milled hardwood material with 10 wt.% bark (to represent forest residues). The results indicate that the presence of hemicellulose and lignin influences the flowability of lignocellulosic feedstocks. Crumbled wood particles show poor flowability at the pilot scale, while the presence of fine materials (less than 0.7 mm) and bark improves the flowability of biomass slurries without affecting the organic acid yields (based on C6 carbohydrate content).

09 - BIOMASS FUELS↗

Exploring Ion Mobility Mass Spectrometry Data File Conversions to Leverage Existing Tools and Enable New Workflows

Ion mobility (IM) is often combined with LC-MS experiments to provide an additional dimension of separation for complex sample analysis. While highly complex samples are better characterized by the full dimensionality of LC-IM-MS experiments to uncover new information, downstream data analysis workflows are often not equipped to properly mine the additional IM dimension. For many samples the data acquisition benefits of including IM separations are all that is necessary to uncover sample information and the full dimensionality of the data is not required for data analysis. Post-acquisition reduction and adaptation of the dimensions of LC-IM-MS and IM-MS experiments into an LC-MS format opens the possibility to use a plethora of existing software tools. In this work, we developed data file conversion tools to reduce the complexity of IM data analysis. Three data file transformations are introduced in the PNNL PreProcessor software: 1) mapping the IM axis to the LC axis for IM-MS data, 2) converting the drift time vs. m/z space to CCS/z vs m/z space, and 3) transforming All Ions IM/MS mobility aligned fragmentation data to a standard LC-MS DDA data file format. Finally, these new data file conversions are demonstrated with corresponding lipidomics and proteomics workflows that leverage existing LC-MS data analysis software to highlight the benefits of the data transformations.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗