Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “preprocessed”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Spatially Accelerated Winding Numbers for Curved Geometry

The generalized winding number (GWN) is a scalar field that supports robust containment queries on curved geometry, including non-watertight, overlapping, and nested boundary representations. While queries can be easily parallelized over samples, direct evaluation on parametric curves and surfaces remains costly for large and complex models. Fast, state-of-the-art GWN approaches leverage a spatial index to approximate the GWN, typically coupled with a Taylor expansion which approximates the GWN contribution for far clusters of geometric primitives. However, such methods operate only on discrete inputs such as triangle meshes and point clouds, and would introduce containment errors near boundaries if applied to curved input. We extend support for fast GWN evaluation over arbitrary collections of NURBS curves in 2D and trimmed NURBS patches in 3D via a Bounding Volume Hierarchy that stores efficiently precomputed moment data in the hierarchy nodes. When querying the hierarchy, approximations for far clusters are used alongside direct evaluation for nearby NURBS primitives, achieving sub-linear complexity while preserving the geometric features in the vicinity of the query point. Central to our performance improvements is an adaptive subdivision strategy for NURBS primitives during a preprocessing phase, creating better spatial partitions while retaining the same accuracy for containment decisions as a direct evaluation. We demonstrate the performance and accuracy of our approach across a large collection of 2D and 3D datasets.

Computer science↗

Physical Interpretation of Early Battery Life Prediction Models

Early battery life prediction models are most useful for R&D if they help us understand the early changes in battery electrochemical response that correspond with long-term degradation and failure. Linear regression models such as Fused lasso and Partial Least Squares can fit coefficients directly to high-dimensional electrochemical data like capacity-voltage and ΔV–state-of-charge, i.e., Q(V) and ΔV(SOC) curves, learning coefficients that can be physically interpreted. We leverage the ISU-ILCC battery aging data set to learn high-dimensional coefficients for early battery life prediction from traditional slow-rate capacity check data, demonstrating learning on Q(V), d Q· d V −1 , and ΔV(SOC) curves. A thorough study on the dependence of coefficient values on train/test size and data preprocessing methods is made, demonstrating the reliability of high-dimensional regression approaches unless very small amounts of data are used for model training. For this data set, coefficients from Q(V) and d Q· d V −1 models highlight changes in electrode stoichiometry due to lithium loss, while ΔV(SOC) coefficients highlight changes in positive electrode diffusivity due to particle cracking as well as electrode stoichiometry shifts. By directly interpreting the coefficients of a regression model, we make physical insights into battery degradation mechanisms without requiring the assumptions of traditional battery data analysis methods.

25 ENERGY STORAGE↗

Diffraction imaging of light induced dynamics in xenon-doped helium nanodroplets

Time-resolved wide-angle coherent diffraction imaging of individual helium nanodroplets, doped with xenon and excited with an 800 nm NIR laser pulse. Raw and preprocessed pattern, together with useful metadata, are contained in HDF5 files. Every field in the H5 container comes with a description, under the field "description", and the actual data under the field "data". For further information about the dataset see the [ArXiv article](https://arxiv.org/abs/2205.04154 ) and Bruno Langbehn's [PhD thesis](https://depositonce.tu-berlin.de/handle/11303/13189).

FERMI FEL-1↗

Diffraction imaging of light induced dynamics in xenon-doped helium nanodroplets

Time-resolved wide-angle coherent diffraction imaging of individual helium nanodroplets, doped with xenon and excited with an 800 nm NIR laser pulse. Raw and preprocessed pattern, together with useful metadata, are contained in HDF5 files. Every field in the H5 container comes with a description, under the field "description", and the actual data under the field "data". For further information about the dataset see the [ArXiv article](https://arxiv.org/abs/2205.04154 ) and Bruno Langbehn's [PhD thesis](https://depositonce.tu-berlin.de/handle/11303/13189).

FERMI FEL-1↗

Moisture Effect on Chitin Decomposition Biogeochemistry

This dataset contains data files for multiple measurements of sample biogeochemistry and function collected for the Soils SFA Chitin Decomposition project in task 2.2. Samples were generated from soil incubated under different moisture levels, with and without chitin. Each sheet in the file refers to the preprocessed data collected. Sheet 1 "Respiration" measures CO2 production daily for the course of the incubation. Sheet 2 "Biomass" contains the microbial biomass and salt extractable measurements for carbon and nitrogen. Sheet 3 "Chitin" is for HPLC measured chitin from each sample. Sheet 4 "Extracellular Enzyme Assays" records the level of activity for several enzyme assays. Sheet 5 "Enzyme Kinetics" measures degradation of substrate over time for all samples.

Reichart, Nicholas J [Pacific Northwest National L↗

Serial2Parallel

In the era of machine learning, we often need to run the same code/script many times with little or no variations (e.g., performance evaluation, data preprocessing, data generation, hyperparameter tuning, etc.). It is not a problem when you just need to do that a few times, but when the number of repetitions becomes very large, it can be a daunting task. The code “Serial2Parallel” provides an easy way for users to be able to run many numbers of any serial code/scripts in a parallel manner across multiple nodes in an message passing interface (MPI) cluster. The code includes the server program that deals with task pool management and client program that processes task. The server gets the tasks ready and waits for clients' connections. The client code pulls tasks from the server and processes them. The client code will run in parallel.

Sangkeun, MattLee↗

Sandia Injury Biomechanics Laboratory Library

The Sandia Injury Biomechanics Laboratory (SIBL) Library is a collection of scripts and tools used to preprocess, process, and postprocess data for analysis of injury biomechanics data. SAND2020-3663 M Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Hovey, Chad↗

Stochastic Techno Economic Model

The Stochastic Techno-Economic Model or STEM is an analytical tool that estimates the logistics cost of different biomass feedstocks by incorporating uncertainty into the modeling framework. The scope of the model covers multiple stages of the biomass life cycle spanning feedstock harvest, collection, transportation and handling, preprocessing, and storage. The model determines the total logistics cost per dry metric ton ($/DM ton) for biomass and breaks down the costs for important cost categories including ownership related costs such as interest and depreciation, insurance, housing, and taxes as well as operating costs like repairs and maintenance, labor costs, and fuel and lube costs.

Burli, PralhadH↗

gp_blendclass_singleband

The code used for the data preprocessing, image simulation, and model training and analysis reported in the paper "Gaussian Process Classification for Galaxy Blend Identification in LSST" (arXiv:2107.09246).

Buchanan, JamesJ.↗

AdversarialTensors

This library builds a framework for defending ML models against adversarial attacks. The library will be developed at various stages leading to publication and software release at each stage. We employ tensor decomposition strategies as preprocessing stages for the first stage to provide robustness against the prominent adversarial noise. In the second stage, we develop a latent noise generator capable of generating novel adversarial noise that threatens the existing state-of-the-art defense strategy. In the third stage, we develop a UNSUP-GAN model, where the generator is trained to denoise against latent noise and most adversarial noises. This generator can provide a robust adversarial attack against any unseen attack.

Bhattarai, Manish↗

Digital Analytics, Causal Knowledge Acquisition and Reasoning for Technical Language Processing

Complex engineering systems such as nuclear power plants (NPPs) generate and collect large amounts of equipment reliability (ER) data elements that contain information on the status of components, assets, and systems. Some of this information is textual in form and can be found in documents such as incident reports (IRs) and work orders (WOs). Analyses of textual data in current NPPs-using natural language processing (NLP) methods-have been expanded over the last decade, and it is only recently that the true potential of such analyses has emerged. So far, applications of NLP methods have mostly been limited to classification and prediction, the goal being to identify the nature of the textual element (e.g., safety or non-safety related). Here, we target a more complex problem: automatically extracting knowledge from a textual element in order to assist system engineers in conducting system health assessments. Knowledge extraction is a very broad concept, and its definition may vary depending on the application context. Our methods are a blend of both rule-based and machine learning (ML) algorithms. For our purposes, knowledge extraction means identifying the systems or assets mentioned in a given textual element, as well as the type of event described (e.g., component failure or maintenance activity). In addition, we want to capture details such as measured quantities and the temporal/cause-effect relations between events. In this tool, we also demonstrate how textual data elements are preprocessed in order to handle typos, acronyms, and abbreviations. One main feature of these methods is that they are not based solely on data, but are in fact model-based. In other words, they also rely on MBSE models that are designed to capture-from a functional point of view-the architecture of the systems/assets under consideration. The main purpose of such models is to digitally emulate system engineers' knowledge of system and asset architecture and to identify dependencies among systems, assets, and components. Provided these models, analyses of textual and numeric ER data can be performed by first identifying the OPM model elements to which the ER data elements are referring. The relationships between ER data elements are then identified by checking for any temporal or logical dependencies.

Mandelli, Diego [Idaho National Laboratory (INL), ↗

Offshore Wind ENergy Simulation Toolkit (OWENS)

SAND2021-2751 O The Offshore Wind ENergy Simulation Toolkit (OWENS) is a collection of aerodynamic, structural, hydrodynamic, drivetrain, controls, composite structure and mesh preprocessing, and data postprocessing. OWENS is primarily an ontology, or glue code, pulling together many open-source and Sandia-developed libraries to model the aero-servo-hydro-elastic physics of wind and marine energy turbines. The toolkit’s intended use is for arbitrary aeroelastic rotor configurations analysis including vertical-axis wind turbines, horizontal-axis wind turbines, and analogous marine energy applications for fixed-bottom and floating configurations. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525

Owens, Brian↗

HydraGNN v4.0

The new version of HydraGNN v4.0 provides additional core capabilities, such as: Inclusion of multi-body atomistic cluster expansion MACE, polarizable atom interaction neural network PAINN, and equivariant principal neighborhood aggregation (PNAEq) among the message passing layers supported -Inclusion of graph transformers to directly model long-range interactions between nodes that are distant in the graph topology Integration of graph transformers with message passing layers by combining the graph embedding generated by the two mechanisms, which allows for an improved expressivity of the HydraGNN architecture Improved re-implementation of multi-task learning (MTL) to allow its use for stabilized training across imbalanced, multi-source, multi-fidelity data Introduction of multi-task parallelism, a newly proposed type of model parallelism specifically for MTL architectures, which allows to dispatch different output decoding heads to different GPU devices Integration of multi-task parallelism with pre-existing distributed data parallelism to enable a 2D parallelization for distributed training Improved portability of the distributed training across Intel GPUs, which has been testes on ALCF exascale supercomputer Aurora Inclusion of 2-level fine-grained energy profilers portable across NVIDIA, AMD, and Intel GPUs to monitor the power and energy consumption associated with different functions executed by the HydraGNN code during data pre-load and training Restructuring of previous examples and inclusion of new sets of examples to illustrate the download, preprocess, and training of HydraGNN models on new large-scale open-source datasets for atomistic materials modeling (e.g., Alexandria, Transition1x, OMat24, OMol25)

Lupo Pasini, Massimiliano [Oak Ridge National Labo↗

COBRA:COMPUTED-TOMOGRAPHY BASED RANDOM-FIELD APPROXIMATION

SF-25-115 COBRA (COmputed-tomography Based Random-field Approximation) is a Python application for generating statistically equivalent random fields from CT-scan imagery. It leverages Karhunen–Loève expansions to model microstructural variability, enabling users to: Preprocess CT scans (filtering and Gaussian transformation); Fit covariance kernels fromempirical data; Solve eigenproblems to obtain KL modes; Sample random fields onsistent with fitted statistics; Postprocess samples back into the physical domain.

Hu, Tianchen↗

Detecting Living-off-the-land Attacks Using K-means And Graph Convolutional Networks

The code ingests Zeek logs derived from network packet captures and goes through data preprocessing before it gets passed into a K-Means model that labels each device as either a client or server. Graph Convolutional Network (GCN) model is used to obtain the embeddings to represent the features in lower dimension. Last, K-means cluster analysis is used to cluster the embeddings for each class.

Quach, Anna [Idaho National Laboratory (INL), Idah↗

Knowledge Oriented Graph Unified Transformer (KOGUT) v0.1

KOGUT — Knowledge Oriented Graph Unified Transformer KOGUT implements the Relational Graph Transformer (RelGT) architecture for knowledge graph link prediction in biological domains, with a primary focus on microbial growth media prediction. While the original RelGT (arXiv:2505.10960) targets relational tables, time series, and multi-table databases, KOGUT adapts this architecture for heterogeneous biological knowledge graphs, providing first-in-class AI predictive models for microbial cultivation. Key Adaptations Beyond Original RelGT: - Knowledge Graph Focus: Applied to biological KGs with semantic node types (taxa, chemicals, media, phenotypes, environments) versus generic relational database tables, trained on the KG-Microbe knowledge graph (1.3M entities, 2.9M edges, 24 relation types). - Multimodal Node Encoding: Integrates node labels, categories, descriptions, and synonyms from KG metadata through learned embedding layers—adapting relational column features to graph node attributes with textual semantics. - Extended K-Hop Subgraph Strategy: Optimized neighborhood sampling (3-hop default, configurable up to 200 nodes) tuned for sparse biological networks, building on the original local-global attention framework with biological relation preservation. - Biolink Predicate Preservation: Type-specific transformations for 24 biological edge semantics (occurs_in, consumes, produces, has_phenotype, subclass_of) beyond standard relational foreign keys, enabling multi-relation link prediction. - Inductive Learning Support: Enables zero-shot predictions for novel taxa through feature-based embeddings (temperature, oxygen requirements, gram stain, cell shape), extending the original transductive relational benchmark scope to uncultured microorganisms. CheapSOTA Performance Optimizations (This Distribution): - VQ-EMA Centroid Attention: Vector quantization with exponential moving average for improved global context modeling (+5-10% MRR improvement). - HDF5 Precomputed Data Loading: One-time preprocessing of k-hop subgraphs to eliminate redundant graph traversals (2-5× training speedup). - Distributed Data Parallel Training: Multi-GPU support for scaling to larger knowledge graphs (tested on 4× NVIDIA A100 GPUs at NERSC Perlmutter). - Mixed Precision Training: Automatic mixed precision (AMP) for memory efficiency and faster training. Advantages Over Standard Knowledge Graph Embedding Models: Combines RelGT's proven multi-element tokenization (features, type, hop, structure) with graph-native biological representations, enabling interpretable link prediction across heterogeneous entities that standard embedding models (TransE, RotatE, ComplEx) and table-based transformers cannot directly model. Achieves near-perfect performance on microbial growth media prediction (MRR: 0.9966, Precision@1: 0.9932, Hit@10: 1.0000) while maintaining explainability through attention-based reasoning over biological pathways. Training Data: - KG-Microbe merged knowledge graph: 1,379,337 nodes, 2,960,472 edges - 24 biological relation types including taxonomic hierarchies, metabolic interactions, phenotype associations, and environmental relationships - Primary prediction task: Growth media suitability for microbial taxa (biolink:occurs_in, 50K edges) - Multi-relation capability: Predicts links for any of the 24 relation types, including chemical consumption/production, phenotype associations, and taxonomic classification Citation: Original RelGT Architecture: Dwivedi et al., "Relational Graph Transformer", arXiv:2505.10960, 2025 KOGUT Implementation: Knowledge Oriented Graph Unified Transformer for Microbial Growth Media Prediction Developed at Lawrence Berkeley National Laboratory (LBNL) Trained on NERSC Perlmutter supercomputer

Joachimiak, Marcin [Lawrence Berkeley National Lab↗

Mass Spectrometer Transient Analysis

This software implements a complete preprocessing pipeline for transient mass spectrometry (MS) data collected during TAP (Temporal Analysis of Products) experiments. It is designed to extract chemically meaningful fluxes from overlapping ion signals by applying a calibrated defragmentation matrix and solving the resulting linear system using non-negative least squares (NNLS) regression. The core script, preprocess_mass_spec.py, performs the following operations: Gain correction: Applies amplifier gain scalars derived from inert-packed calibration pulses to normalize signal intensities across AMUs and acquisition settings. Background subtraction: Removes experiment baselines using user-defined time windows, ensuring compatibility with slow-diffusing species and preventing negative values that would interfere with NNLS. Options to subtract before and after defragmentation. Defragmentation: Constructs a fragmentation matrix A from zeroth moments of calibration pulses (equal molar gas:inert mixtures) and solves Ax=b at each time point, where b is the raw MS signal and x is the estimated species flux. The matrix is normalized to inert signals and accounts for instrument-specific fragmentation behavior. Pulse-mode handling: Supports both averaged and individual pulse modes, enabling statistical treatment of fluxes and calculation of standard deviations. Integration and output: Computes zeroth moments (integrated fluxes) and exports time-resolved and integrated data in CSV format, suitable for downstream kinetic modeling. The software is validated using both virtual TAP simulations (VTAP) and experimental data from propane dehydrogenation (PDH) on CrOx/Al2O3 catalysts. It preserves temporal resolution by applying NNLS point-by-point across the pulse duration (typically 6,000+ time slices per pulse), leveraging the linear superposition principle to reconstruct full flux profiles. The defragmented outputs are compatible with kinetic extraction methods such as the G and Y procedures, which are used to derive rate–concentration relationships from TAP data. The details of these validations are discussed in detail in the supporting manuscript and supporting information. Example data and output files are also included. The methodology is robust to experimental noise and drift, with calibration protocols that account for pulse size effects, MS aging, and inert gas normalization. The software is modular, reproducible, and tailored for high-throughput TAP-MS workflows in catalysis research.

Kristy, Stephen [Idaho National Laboratory (INL), ↗

Framework For Spatial Agricultural Crop Yield Prediction Model Development

This framework was developed to provide data preprocessing for spatiotemporal agricultural yield data and remote sensing data for modelling using artificial neural networks (ANNs) to predict subfield crop yield estimates. The software includes methods to train, validate, and test ANN models. It also include methods to infer on new remote sensing data.

Griffel, LloydM.↗