Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “decoding”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19

Gradient Coding With Iterative Block Leverage Score Sampling

Gradient coding is a method for mitigating straggling servers in a centralized computing network that uses erasure-coding techniques to distributively carry out first-order optimization methods. Randomized numerical linear algebra uses randomization to develop improved algorithms for large-scale linear algebra computations. In this study, we propose a method for distributed optimization that combines gradient coding and randomized numerical linear algebra. The proposed method uses a randomized ℓ 2 -subspace embedding and a gradient coding technique to distribute blocks of data to the computational nodes of a centralized network, and at each iteration the central server only requires a small number of computations to obtain the steepest descent update. The novelty of our approach is that the data is replicated according to importance scores, called block leverage scores, in contrast to most gradient coding approaches that uniformly replicate the data blocks. Furthermore, we do not require a decoding step at each iteration, avoiding a bottleneck in previous gradient coding schemes. We show that our approach results in a valid ℓ 2 -subspace embedding, and that our resulting approximation converges to the optimal solution.

97 MATHEMATICS AND COMPUTING↗

Performance Optimization for Relative-Error-Bounded Lossy Compression on Scientific Data

Scientific simulations in high-performance computing (HPC) environments generate vast volume of data, which may cause a severe I/O bottleneck at runtime and a huge burden on storage space for postanalysis. Unlike traditional data reduction schemes such as deduplication or lossless compression, not only can error-controlled lossy compression significantly reduce the data size but it also holds the promise to satisfy user demand on error control. Pointwise relative error bounds (i.e., compression errors depends on the data values) are widely used by many scientific applications with lossy compression since error control can adapt to the error bound in the dataset automatically. Pointwise relative-error-bounded compression is complicated and time consuming. In this article, we develop efficient precomputation-based mechanisms based on the SZ lossy compression framework. Our mechanisms can avoid costly logarithmic transformation and identify quantization factor values via a fast table lookup, greatly accelerating the relative-error-bounded compression with excellent compression ratios. In addition, we reduce traversing operations for Huffman decoding, significantly accelerating the decompression process in SZ. Experiments with eight well-known real-world scientific simulation datasets show that our solution can improve the compression and decompression rates (i.e., the speed) by about 40 and 80 p, respectively, in most of cases, making our designed lossy compression strategy the best-in-class solution in most cases.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Progressive Tree-Based Compression of Large-Scale Particle Data

Scientific simulations and observations using particles have been creating large datasets that require effective and efficient data reduction to store, transfer, and analyze. However, current approaches either compress only small data well while being inefficient for large data, or handle large data but with insufficient compression. Toward effective and scalable compression/decompression of particle positions, we introduce new kinds of particle hierarchies and corresponding traversal orders that quickly reduce reconstruction error while being fast and low in memory footprint. Our solution to compression of large-scale particle data is a flexible block-based hierarchy that supports progressive, random-access, and error-driven decoding, where error estimation heuristics can be supplied by the user. For low-level node encoding, we introduce new schemes that effectively compress both uniform and densely structured particle distributions. Our proposed methods thus target all three phases of a tree-based particle compression pipeline, namely tree construction, tree traversal, and node encoding. In conclusion, the improved efficacy and flexibility of these methods over existing compressors are demonstrated through extensive experimentation, using a wide range of scientific particle datasets.

97 MATHEMATICS AND COMPUTING↗

Non-Blind Deblurring for Fluorescence: A Deformable Latent Space Approach with Kernel Parameterization

We report N\non-blind deblurring (NBD) is a modeling method of the image deblurring problem in computer vision, where the blurring kernel is known or can be externally estimated. In this paper, we attempt to solve a parametric NBD problem, inspired by the simultaneous acquisition of ptychography and fluorescent imaging (FI). Ptychography is an imaging method that favors larger probes, i.e. convolutional kernels, while FI relies on a small probe for high resolution. Also, the kernel can be solved during ptychographic reconstruction. With Ptycho-FI using the same larger kernel, we can perform NBD on the blurred fluorescent images to achieve high-resolution FI, and thus speed up the experiments. To this end, we design a deep latent space deformation network that is directly parameterized by the kernel. The network consists of three components: encoder, deformer, and decoder, where the deformer is specifically meant to rectify the latent space representations of blurred images to a standard latent space, regardless of the kernel. The deformation network is trained with a two-stage training scheme. We conduct extensive experiments to confirm that our parametric model can adapt to drastically different blurring kernels and perform robust deblurring.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

Deep learning multiphysics network for imaging CO 2 saturation and estimating uncertainty in geological carbon storage

Multiphysics inversion exploits different types of geophysical data that often complement each other and aims to improve overall imaging resolution and reduce uncertainties in geophysical interpretation. Despite the advantages, traditional multiphysics inversion is challenging because it requires a large amount of computational time and intensive human interactions for preprocessing data and finding trade-off parameters. These issues make it nearly impossible for traditional multiphysics inversion to be applied as a real-time monitoring tool for geological carbon storage. In this paper, we present a deep learning (DL) multiphysics network for imaging CO 2 saturation in real time. The multiphysics network consists of three encoders for analysing seismic, electromagnetic and gravity data and shares one decoder for combining imaging capabilities of the different geophysical data for better predicting CO 2 saturation. The network is trained on pairs of CO 2 label models and multiphysics data so that it can directly image CO 2 saturation. Here we use the bootstrap aggregating method to enhance the imaging accuracy and estimate uncertainties associated with CO 2 saturation images. Using realistic CO 2 label models and multiphysics data derived from the Kimberlina CO 2 storage model, we evaluate the performance of the deep learning multiphysics network and compare its imaging results to those from the deep learning single-physics networks. Our modelling experiments show that the deep learning multiphysics network for seismic, electromagnetic, and gravity data not only improves the imaging accuracy but also reduces uncertainties associated with CO 2 saturation images. Our results also suggest that the deep learning multiphysics network for the non-seismic data (i.e., electromagnetic and gravity) can be used as an effective low-cost monitoring tool in between regular seismic monitoring.

58 GEOSCIENCES↗

An automated workflow that generates atom mappings for large‐scale metabolic models and its application to Arabidopsis thaliana

SUMMARY Quantification of reaction fluxes of metabolic networks can help us understand how the integration of different metabolic pathways determines cellular functions. Yet, intracellular fluxes cannot be measured directly but are estimated with metabolic flux analysis (MFA), which relies on the patterns of isotope labeling of metabolites in the network. The application of MFA also requires a stoichiometric model with atom mappings that are currently not available for the majority of large‐scale metabolic network models, particularly of plants. While automated approaches such as the Reaction Decoder Toolkit (RDT) can produce atom mappings for individual reactions, tracing the flow of individual atoms of the entire reactions across a metabolic model remains challenging. Here we establish an automated workflow to obtain reliable atom mappings for large‐scale metabolic models by refining the outcome of RDT, and apply the workflow to metabolic models of Arabidopsis thaliana . We demonstrate the accuracy of RDT through a comparative analysis with atom mappings from a large database of biochemical reactions, MetaCyc. We further show the utility of our automated workflow by simulating 15 N isotope enrichment and identifying nitrogen (N)‐containing metabolites which show enrichment patterns that are informative for flux estimation in future 15 N‐MFA studies of A. thaliana . The automated workflow established in this study can be readily expanded to other species for which metabolic models have been established and the resulting atom mappings will facilitate MFA and graph‐theoretic structural analyses with large‐scale metabolic networks.

59 BASIC BIOLOGICAL SCIENCES↗

A New Evaluation Metric for Demand Response-Driven Real-Time Price Prediction Towards Sustainable Manufacturing

Abstract The increasing industry energy demand highlights the urgency of demand response management, while the emerging smart manufacturing technologies pave the way for the implementation of real-time price (RTP)-based demand response management towards sustainable manufacturing. The demand response management requires scheduling of manufacturing systems based on RTP predictions, and thus the prediction quality can directly alter the effectiveness of demand response. However, since the general price prediction algorithms and prediction evaluation metrics are not specifically designed for RTP in demand response problems, a good RTP prediction obtained and evaluated by these algorithms and metrics may not be suitable for demand response scheduling. Therefore, in this study, the relationships between the effectiveness of demand response for manufacturing systems and evaluation results from six commonly used metrics are investigated. Meanwhile, a new metric called k-peak distance (KPD), considering the characteristics of the demand response problem, is proposed and compared with the other six metrics. Furthermore, an encoder-decoder long short-term memory recurrent neural network with KPD is proposed to provide better RTP prediction for manufacturing demand response problems. The case studies indicate that the proposed KPD metric shows a 1.8–3.6 times higher correlation with the demand response effectiveness compared to the other metrics. In addition, the production schedule based on the RTP prediction obtained from the proposed algorithm can improve the effectiveness of demand response by 23.4% on average.

Engineering↗

A Variational Autoencoder Model Toward Molecular Structure Representation Learning of Fuels

Here, in this work, a Variational Autoencoder (VAE)-based data-driven modeling framework is developed with the overarching goal of enabling fuel design. The VAE model is trained on a large dataset with several chemical species to learn a compressed latent space molecular representation. Chemical structure in the form of Simplified Molecular Input Line Entry System (SMILES) string is fed as input, encoded into the VAE latent space, and decoded back to the SMILES string using Long Short-Term Memory (LSTM) networks. Complexities of the VAE training loss function are thoroughly examined by varying the weightage (beta (𝜷) parameter) of the latent space regularization term, thereby assessing the balance between reconstruction accuracy and validity, and focusing on both accurate molecular structure reconstruction and latent space consistency. Two different strategies for 𝜷 variation are evaluated: linear annealing and cyclic annealing. In addition, the impact of total correlation adjustment and hierarchical priors is also studied with regard to the balance between reconstruction fidelity and latent space regularization, and potential issues such as posterior collapse, over-regularization, and poor disentanglement of latent variables. Overall, the best performance of the model is achieved with hierarchical priors and incrementally increasing 𝜷 from 0 to a threshold value of 0.25 over 75 epochs. The generative VAE model can be readily coupled with Quantitative Structure–Property Relationship (QSPR) analysis to develop an integrated end-to-end framework for fuel-property prediction and molecular design of novel promising fuels.

fuel design↗

Regulation of translation by ribosomal RNA pseudouridylation

Pseudouridine is enriched in ribosomal, spliceosomal, transfer, and messenger RNA and thus integral to the central dogma. The chemical basis for how pseudouridine affects the molecular apparatus such as ribosome, however, remains elusive owing to the lack of structures without this natural modification. Here, we studied the translation of a hypopseudouridylated ribosome initiated by the internal ribosome entry site (IRES) elements. We analyzed eight cryo–electron microscopy structures of the ribosome bound with the Taura syndrome virus IRES in multiple functional states. We found widespread loss of pseudouridine-mediated interactions through water and long-range base pairings. In the presence of the translocase, eukaryotic elongation factor 2, and guanosine 5'-triphosphate hydrolysis, the hypopseudouridylated ribosome favors a rare unconducive conformation for decoding that is partially recouped in the ribosome population that remains modified at the P-site uridine. The structural principles learned establish the link between functional defects and modification loss and are likely applicable to other pseudouridine-associated processes.

59 BASIC BIOLOGICAL SCIENCES↗

Analysis of a Runtime Data Sharing Architecture over LTE for a Heterogeneous CAV Fleet

This paper describes a lightweight runtime architecture for telemetry, communication, and control of cars deployed with advanced driver assistance systems where a human is in the loop with the car, via an LTE connection. The system architecture supports both local control decisions based on car sensors and safety algorithms as well as high-level input from external systems that may provide insight into traffic state ahead of sensor data. Implementation of the architecture is done in ROS and depends on open-source software packages for runtime decoding of information from the vehicle’s controller area network (CAN) and integration of GPS data from accompanying sensors. The contribution of the paper is to describe the overall architecture, the data it can communicate to other systems, performance of the system at runtime, and challenges faced when deploying the architecture across a heterogeneous fleet. Preliminary results from analysis of test data will provide insights into whether the use of high-latency communication can be effective for societal-scale intelligent transportation systems when applied in future scenarios

Richardson, Alex↗

Demonstrating Cross-Facility Data Processing at Scale with Laue Microdiffraction

In February and April 2023 live, at-scale data processing demonstrations were conducted between the Advanced Photon Source (APS), a synchrotron light source, and the Argonne Leadership Computing Facility (ALCF). These tests were run as part of a novel beamline technique: coded aperture laue micro-diffraction. This technique requires a significant amount of compute to decode appeture patterns embedded in the detector stream. An autonomous system was able to send data to ALCF during an experiment, utilize 50 nodes of the Polaris supercomputer to process 6-12 hour scans, and return the data back to the APS within 12-15 minutes behind the detector. With scan points arriving every 72 seconds, the system kept up with the beamline, potentially enabling in-experiment analysis. The data processing system utilizes Globus infrastructure and an on-demand queue to dynamically acquire nodes on Polaris. The underlying reconstruction algorithms were parallelized via MPI and accelerated with custom CUDA kernels.

Prince, Michael↗

Fast 2D Bicephalous Convolutional Autoencoder for Compressing 3D Time Projection Chamber Data

High-energy large-scale particle colliders produce data at high speed in the order of 1 terabytes per second in nuclear physics and petabytes per second in high energy physics. Developing real-time data compression algorithms to reduce such data at high throughput to fit permanent storage has drawn increasing attention. Specifically, at the newly constructed sPHENIX experiment at the Relativistic Heavy Ion Collider (RHIC), a time projection chamber is used as the main tracking detector, which records particle trajectories in a volume of three-dimensional (3D) cylinder. The resulting data are usually very sparse with occupancy around 10.8%. Such sparsity presents a challenge to conventional learning-free lossy compression algorithms, such as SZ, ZFP, and MGARD. The 3D convolutional neural network (CNN)-based approach, Bicephalous Convolutional Autoencoder (BCAE), outperforms traditional methods both in compression rate and reconstruction accuracy. BCAE can also utilize the computation power of graphical processing units suitable for deployment in a modern heterogeneous highperformance computing environment. This work introduces two BCAE variants: BCAE++ and BCAE-2D. BCAE++ achieves a 15% better compression ratio and a 77% better reconstruction accuracy measured in mean absolute error compared with BCAE. BCAE-2D treats the radial direction as the channel dimension of an image, resulting in a 3× speedup in compression throughput. In addition, we demonstrate an unbalanced autoencoder with a larger decoder can improve reconstruction accuracy without significantly sacrificing throughput. Lastly, we observe both the BCAE++ and BCAE-2D can benefit more from using half-precision mode in throughput (76 - 79% increase) without loss in reconstruction accuracy. The source code and links to data and pretrained models can be found at https://github.com/BNL-DAQ-LDRD/NeuralCompression_v2

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Emerging Flexible Designs for Geospatial Multimodal Foundation Models

Foundation models are rapidly transforming Earth observation by enabling scalable pretraining across diverse unlabeled geospatial modalities. However, their architectural diversity—ranging from encoder-only to encoder-decoder and masked autoencoding paradigms—makes it challenging to assess performance trade-offs in a consistent manner. In this work, we present an apples-to-apples comparison of leading FM architectures designed for geospatial multimodal reasoning, with a particular focus on flexibility across varied spectral band configurations. We standardize pretraining using identical self-supervised learning objectives and training datasets, and evaluate all models under consistent parameterization on the GEOBench benchmark across classification and segmentation tasks. Our results offer new insights into the design trade-offs between model flexibility, modality alignment, and downstream task performance. By highlighting architectural strengths and limitations under controlled conditions, this study provides practical guidance for building next-generation geospatial foundation models capable of robust multimodal reasoning.

Ambrozio Dias, Philipe [ORNL] (ORCID:0000000194277↗

Combined Mixed Potential Electrochemical Sensors and Artificial Neural Networks for the Quantificationand Identification of Methane in Natural Gas Emissions Monitoring

Sensors capable of quantifying methane concentration and discriminating between possible sources are needed for natural gas leak detection where multiple spatially overlapping sources including wetlands and agriculture may be present. We report on the fabrication by an additive manufacturing process of a four electrode La 0.87 Sr 0.13 CrO 3 , Indium Tin Oxide (In 2 O 3 90 wt%, SnO 2 10 wt%), Au, Pt mixed potential electrochemical sensor using yttria-stabilized zirconia (YSZ) as a solid electrolyte to natural gas detection. Artificial neural networks (ANNs) are used to automatically decode the possible source and concentration of methane. The ANNs trained on sensor data are capable of correctly discriminating between three sources of methane emissions from simulated mixtures of emissions from cattle, wetlands, or natural gas with >98% accuracy. Quantification error for methane in mixtures of CH 4 in air, CH 4 + NH3 in air, and simulated natural gas is less than 1.5% ppm when a two-temperature dataset is employed.

03 NATURAL GAS↗

InterSpec v. 1.0.8

SAND2021-4823 O InterSpec decodes and displays gamma radiation data from a variety of handheld, laboratory, and fixed installation detector types. Once loaded, InterSpec provides a means to interactively analyze the data via a peak-based methodology. InterSpec assists with isotope identification and determines the source strength and shielding characteristics of the measured item. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Johnson, William↗

labquake_future_prediction

The labquake_future_prediction code is a collection of python modules and scripts that serves as supporting information for the article “Predicting future laboratory fault friction through deep learning” for publication in the journal of “Geophysical Research Letters”. It is designed to predict laboratory fault slips in the immediate future by scanning continuous acoustic emission (AE) waveforms recorded in laboratory biaxial shear experiments. The predictions are made with a deep learning model based on convolutional encoder-decoder (CED) models and the Transformer model primarily developed for Natural Language Processing (NLP). The deep learning model is trained with the tensorflow package using publicly available laboratory data sets in standard binary file format in numpy. The utility functions for reading data files, configuring model hyperparameters, constructing the CED and Transformer models, training and testing of the models are defined in python module files. The workflow of training the models for labquake future predictions and the multiple GPU’s rapid model hyperparameter optimization as described in the journal article, are demonstrated in accompanying python script files and Jupyter notebooks.

Wang, Kun↗

InversionNet

InversionNet is a software to solve subsurface imaging problems. It leverages a convolutional neural network with an encoder-decoder structure to model the correspondence from seismic data to subsurface velocity structures.

Lin, Youzuo↗

InterSpec v. 1.0.9

SAND2021-4823 O InterSpec decodes and displays gamma radiation data from a variety of handheld, laboratory, and fixed installation detector types. Once loaded, InterSpec provides a means to interactively analyze the data via a peak-based methodology. InterSpec assists with isotope identification and determines the source strength and shielding characteristics of the measured item. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Johnson, William↗