Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Reduction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Developing Methodology to Determine Pu Isotopic Composition by Laser Ablation MC-ICP-MS

This project will develop methodology to analyze the Pu isotope ratio in mixed U-Pu particles by laser ablation MC-ICP-MS. This will involve: 1) testing and validation of the Pu analytical method using mixed U-Pu solutions and Pu doped glasses, 2) isotopic analysis of mixed U-Pu particles by laser ablation MC-ICP-MS, and 3) continued development of the R-based data reduction program. Details regarding the analytical method development and results of QC testing will be output as a deliverable to the IAEA, along with an updated version of the LARA data reduction software.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Understanding and Estimating Error Propagation in Neural Networks for Scientific Data Analysis

Neural networks are increasingly integrated into scientific discovery, where input data reduction and model quantization play a key role in accelerating inference. However, understanding and mitigating the impact of these techniques on output error is critical for ensuring reliable results, particularly in tasks demanding high numerical precision. This paper introduces a comprehensive framework for optimizing neural network inference in scientific computing by combining data reduction and weight quantization while maintaining error-controlled outcomes. We develop theoretical analyses to bound error propagation under these reductions and propose a framework that balances computational performance with error constraints. Evaluation on real-world learning-based combustion simulations and satellite image classification demonstrates that our derived error bounds accurately predict observed errors while enabling significant computational speedup under our framework. This work highlights the potential for further leveraging advancements in modern lossy compression algorithms and hardware accelerators that support lower-precision formats.

He, Weiming [New Jersey Institute of Technology]↗

SonicPy: a suite of programs for ultrasound pulse-echo data acquisition and analysis

Sound speed and elastic constants measurements in solids and liquids are commonly performed using the ultrasound pulse-echo technique. Recent advances have expanded the use of this technique at numerous high pressure synchrotron beamlines and offline laboratories. However, the increased experimental throughput has revealed many limitations in existing software for handling the rapid measurement and the subsequent data-reduction. Here, we report the development of a collection of computer programs for sound speed measurements using the ultrasound pulse-echo technique, compatible with stepped multi-frequency, as well as broadband-pulse, couplant-corrected methods. The programs provide a highly interactive graphical interface, enable efficient measurement, exploration and near real-time analysis of the ultrasound data, and contain features useful for working with samples under high pressure and/or high temperature. The included analysis programs can alleviate the time required for data reduction from hours to less than a minute, allowing users to make timely and informed decisions regarding the appropriate experimental parameters.

97 MATHEMATICS AND COMPUTING↗

Subaru Hyper Suprime-Cam Survey of Cygnus OB2 Complex – I. Introduction, photometry, and source catalogue

ABSTRACT Low-mass star formation inside massive clusters is crucial to understand the effect of cluster environment on processes like circumstellar disc evolution, planet, and brown dwarf formation. The young massive association of Cygnus OB2, with a strong feedback from massive stars, is an ideal target to study the effect of extreme environmental conditions on its extensive low-mass population. We aim to perform deep multiwavelength studies to understand the role of stellar feedback on the IMF, brown dwarf fraction and circumstellar disc properties in the region. We introduce here, the deepest and widest optical photometry of 1.5○ diameter region centred at Cygnus OB2 in r2, i2, z, and Y-filters, using Subaru Hyper Suprime-Cam (HSC). This work presents the data reduction, source catalogue generation, data quality checks, and preliminary results about the pre-main sequence sources. We obtain 713 529 sources in total, with detection down to ∼28, 27, 25.5, and 24.5 mag in r2, i2, z, and Y-band, respectively, which is ∼3 – 5 mag deeper than the existing Pan-STARRS and GTC/OSIRIS photometry. We confirm the presence of a distinct pre-main sequence branch by statistical field subtraction of the central 18 arcmin region. We find the median age of the region as ∼5 ± 2 Myr with an average disc fraction of ∼9 per cent. At this age, combined with A $_V\, \sim$ 6 – 8 mag, we detect sources down to a mass range of ∼0.01–0.17 M⊙. The deep HSC catalogue will serve as the groundwork for further studies on this prominent active young cluster.

Gupta, Saumya (ORCID:0000000161843958)↗

Demonstration of neutron time-of-flight diffraction with an event-mode imaging detector

Neutron diffraction beamlines have traditionally relied on deploying large detector arrays of 3 He tubes or neutron-sensitive scintillators coupled with photomultipliers to efficiently probe crystallographic and microstructure information of a given material. Given the large upfront cost of custom-made data acquisition systems and the recent scarcity of 3 He, new diffraction beamlines or upgrades to existing ones demand innovative approaches. This paper introduces a novel Timepix3-based event-mode imaging neutron diffraction detector system as well as first results of a silicon powder diffraction measurement made at the HIPPO neutron powder diffractometer at the Los Alamos Neutron Science Center. Notably, these initial measurements were conducted simultaneously with the 3 He array on HIPPO, enabling direct comparison. Data reduction for this type of data was implemented in the MAUD code, enabling Rietveld analysis. Results from the Timepix3-based setup and HIPPO were benchmarked against McStas simulations, showing good agreement for peak resolution. With further development, systems such as the one presented here may substantially reduce the cost of detector systems for new neutron instrumentation as well as for upgrades of existing beamlines.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

GTC Follow-up Observations of Very Metal-poor Star Candidates from DESI

Abstract The observations from the Dark Energy Spectroscopic Instrument (DESI) will significantly increase the numbers of known extremely metal-poor stars by a factor of ∼10, improving the sample statistics to study the early chemical evolution of the Milky Way and the nature of the first stars. In this paper we report follow-up observations with high signal-to-noise ratio of nine metal-poor stars identified during the DESI commissioning with the Optical System for Imaging and Low-Resolution Integrated Spectroscopy (OSIRIS) instrument on the 10.4 m Gran Telescopio Canarias. The analysis of the data using a well-vetted methodology confirms the quality of the DESI spectra and the performance of the pipelines developed for the data reduction and analysis of DESI data.

79 ASTRONOMY AND ASTROPHYSICS↗

Maintaining Trust in Reduction: Preserving the Accuracy of Quantities of Interest for Lossy Compression

As the growth of data sizes continues to outpace computational resources, there is a pressing need for data reduction techniques that can significantly reduce the amount of data and quantify the error incurred in compression. Compressing scientific data presents many challenges for reduction techniques since it is often on non-uniform or unstructured meshes, is from a high-dimensional space, and has many Quantities of Interests (QoIs) that need to be preserved. To illustrate these challenges, we focus on data from a large scale fusion code, XGC. XGC uses a Particle-In-Cell (PIC) technique which generates hundreds of PetaBytes (PBs) of data a day, from thousands of timesteps. XGC uses an unstructured mesh, and needs to compute many QoIs from the raw data, f.One critical aspect of the reduction is that we need to ensure that QoIs derived from the data (density, temperature, flux surface averaged momentums, etc.) maintain a relative high accuracy. We show that by compressing XGC data on the high-dimensional, nonuniform grid on which the data is defined, and adaptively quantizing the decomposed coefficients based on the characteristics of the QoIs, the compression ratios at various error tolerances obtained using a multilevel compressor (MGARD) increases more than ten times. We then present how to mathematically guarantee that the accuracy of the QoIs computed from the reduced f is preserved during the compression. We show that the error in the XGC density can be kept under a user-specified tolerance over 1000 timesteps of simulation using the mathematical QoI error control theory of MGARD, whereas traditional error control on the data to be reduced does not guarantee the accuracy of the QoIs.

Gong, Qian↗

Final report- UFL - RAPIDS2: A SciDAC Institute for Computer Science, Data, and Artificial Intelligence

The research initiatives supported by the U.S. Department of Energy (DOE) Grant DE-SC0022265 are fundamentally aimed at pioneering advanced machine learning (ML) techniques for scientific data compression within high-performance computing (HPC) environments. This comprehensive body of work addresses the critical challenge posed by the exponential growth of data generated by scientific simulations in domains such as fusion energy, climate modeling, and computational fluid dynamics (CFD). A core objective is to develop compression algorithms that achieve substantial data reduction—often by orders of magnitude—while rigorously ensuring the fidelity of both the primary data (PD) and scientifically crucial derived quantities of interest (QoI). The methodologies deployed under this grant integrate sophisticated deep learning architectures, prominently featuring autoencoders, advanced generative models like conditional diffusion, and hybrid learning techniques. Key innovations include the development of Guaranteed Autoencoders (GAE) and the Guaranteed Conditional Diffusion with Tensor Correction (GCDTC) framework, which provide explicit, instance-level error bounds on reconstructed data. Furthermore, specialized strategies such as nonlinear constraint satisfaction are employed to preserve the integrity of QoI, a vital requirement for the trustworthiness of downstream scientific analyses. This research also focuses on the design and implementation of scalable, GPU-accelerated software pipelines that seamlessly integrate into existing HPC workflows, ensuring both computational efficiency and practical applicability. The CAESAR framework, for example, unifies foundation and generative models to create an adaptive and efficient compression solution for spatio-temporal scientific data. Collectively, these efforts represent a significant advancement in mitigating the scientific data deluge, enabling more effective data management, accelerated scientific discovery, and optimized utilization of HPC resources.

97 MATHEMATICS AND COMPUTING↗

Online data analysis and reduction: An important co-design motif for extreme-scale computers

A growing disparity between supercomputer computation speeds and I/O rates means that it is rapidly becoming infeasible to analyze supercomputer application output only after that output has been written to a file system. Instead, data-generating applications must run concurrently with data reduction and/or analysis operations, with which they exchange information via high-speed methods such as interprocess communications. The resulting parallel computing motif, online data analysis and reduction (ODAR), has important implications for both application and HPC systems design. Here we introduce the ODAR motif and its co-design concerns, describe a co-design process for identifying and addressing those concerns, present tools that assist in the co-design process, and present case studies to illustrate the use of the process and tools in practical settings.

Data Analysis↗

LCLS Big Data Handling – How I Learned to Stop Worrying and Love the Data Deluge

Advanced data and computing systems are vital to Linac Coherent Light Source (LCLS) operations, data interpretation and overall scientific productivity. The transition to MHz-era operation marks a fundamental change in scale that requires new infrastructure and architectures to link LCLS to the required scale of computing needed for scientific interpretation. The LCLS-II Data System meets big data challenges by implementing configurable data reduction that can adapt to multiple science areas, real-time analysis frameworks to provide visualization and fast feedback, and the ability to transfer data to local and remote computational facilities for near real time analysis at the appropriate scale. Feature extracted information generated in the data analysis pipeline - at the edge, local compute, or remote High-Performance Computing (HPC) resources - can be used to steer experiments and inform user decisions during beam time. Artificial Intelligence and Machine Learning (AI/ML) techniques present new opportunities to rapidly analyse large datasets and direct experiments, but create new challenges in scaling, adaptability, complexity, and trustworthiness. We describe how the LCLS-II Data System architecture addresses its data-driven challenges in the areas of data acquisition, data processing, data management, and workflow orchestration to decrease the overall time-to-science and provide a vision for future developments.

artificial intelligence↗

Cleaning Images with Gaussian Process Regression

Many approaches to astronomical data reduction and analysis cannot tolerate missing data: corrupted pixels must first have their values imputed. This paper presents astrofix, a robust and flexible image imputation algorithm based on Gaussian process regression. Through an optimization process, astrofix chooses and applies a different interpolation kernel to each image, using a training set extracted automatically from that image. It naturally handles clusters of bad pixels and image edges and adapts to various instruments and image types. For bright pixels, the mean absolute error of astrofix is several times smaller than that of median replacement and interpolation by a Gaussian kernel. We demonstrate good performance on both imaging and spectroscopic data, including the SBIG 6303 0.4 m telescope and the FLOYDS spectrograph of Las Cumbres Observatory and the CHARIS integral-field spectrograph on the Subaru Telescope.

42 ENGINEERING↗

Scalable Incremental Checkpointing using GPU-Accelerated De-Duplication

Writing large amounts of data concurrently to stable storage is a typical I/O pattern of many HPC workflows. This pattern introduces high I/O overheads and results in increased storage space utilization especially for workflows that need to capture the evolution of data structures with high frequency as checkpoints. In this context, many applications, such as graph pattern matching, perform sparse updates to large data structures between checkpoints. For these applications, incremental checkpointing techniques that save only the differences from one checkpoint to another can dramatically reduce the checkpoint sizes, I/O bottlenecks, and storage space utilization. However, such techniques are not without challenges: it is non-trivial to transparently determine what data has changed since a previous checkpoint and assemble the differences in a compact fashion that does not result in excessive metadata. State-of-art data reduction techniques (e.g., compression and de-duplication) have significant limitations when applied to modern HPC applications that leverage GPUs: slow at detecting the differences, generate a large amount of metadata to keep track of the differences, and ignore crucial spatiotemporal checkpoint data redundancy. This paper addresses these challenges by proposing a Merkle tree-based incremental checkpointing method to exploit GPUs' high memory bandwidth and massive parallelism. Experimental results at scale show a significant reduction of the I/O overhead and space utilization of checkpointing compared with state-of-the-art incremental checkpointing and compression techniques.

Tan, Nigel↗

Codebase release 2.0 for sauce

Low energy nuclear physics experiments are transitioning towards fully digital data acquisition systems. Realizing the gains in flexibility afforded by these systems relies on equally flexible data reduction techniques. In this paper, methods utilizing data frames and in-memory techniques to work with data, including data from self-triggering, digital data acquisition systems, are discussed within the context of a Python package, sauce. It is shown that data frame operations can encompass common analysis needs and allow interactive data analysis. Two event building techniques, dubbed referenced and referenceless event building, are shown to provide a means to transform raw list mode data into correlated multi-detector events. These techniques are demonstrated in the analysis of two example data sets.

Marshall, Caleb (ORCID:0000000211942920)↗

The ECP ALPINE project: In situ and post hoc visualization infrastructure and analysis capabilities for exascale

A significant challenge on an exascale computer is the speed at which we compute results exceeds by many orders of magnitude the speed at which we save these results. Therefore the Exascale Computing Project (ECP) ALPINE project focuses on providing exascale-ready visualization solutions including in situ processing. In situ visualization and analysis runs as the simulation is run, on simulations results are they are generated avoiding the need to save entire simulations to storage for later analysis. The ALPINE project made post hoc visualization tools, ParaView and VisIt, exascale ready and developed in situ algorithms and infrastructures. The suite of ALPINE algorithms developed under ECP includes novel approaches to enable automated data analysis and visualization to focus on the most important aspects of the simulation. Many of the algorithms also provide data reduction benefits to meet the I/O challenges at exascale. ALPINE developed a new lightweight in situ infrastructure, Ascent.

97 MATHEMATICS AND COMPUTING↗

Information-Theoretic Exploration of Multivariate Time-Varying Image Databases

Modern scientific simulations produce very large datasets, making interactive exploration of such data computationally prohibitive. An increasingly common data reduction technique is to store visualizations and other data extracts in a database. The Cinema project is one such approach, storing visualizations in an image database for post hoc exploration and interactive image-based analysis. This work focuses on developing efficient algorithms that can quantify various types of multivariate dependencies existing within multi-variable datasets. It applies specific mutual information measures for the quantification of salient regions from multivariate image data. Here, using such information measures, the opacity of the images is modulated so that the salient regions are automatically highlighted and the domain scientists can interactively explore the most relevant regions for scientific discovery.

97 MATHEMATICS AND COMPUTING↗

Multiresolution classification of turbulence features in image data through machine learning

During large-scale simulations, intermediate data products such as image databases have become popular due to their low relative storage cost and fast in-situ analysis. Serving as a form of data reduction, these image databases have become more acceptable to perform data analysis on. In this work, we present an image-space detection and classification system for extracting vortices at multiple scales through wavelet-based filtering. A custom image-space descriptor is used to encode a large variety of vortex-types and a machine learning system is trained for fast classification of vortex regions. By combining a radial-based histogram descriptor, a bag of visual words feature descriptor, and a support vector machine, our results show that we are able to detect and classify vortex features at various sizes at multiple scales. Once trained, our framework enables the fast extraction of vortices on new, unknown image datasets for flow analysis.

97 MATHEMATICS AND COMPUTING↗

Workflows Community Summit 2022: A Roadmap Revolution

Scientific workflows have become integral tools in broad scientific computing use cases. Science discovery is increasingly dependent on workflows to orchestrate large and complex scientific experiments that range from the execution of a cloud-based data preprocessing pipeline to multi-facility instrument-to-edge-to-HPC computational workflows. Given the changing landscape of scientific computing (often referred to as a computing continuum) and the evolving needs of emerging scientific applications, it is paramount that the development of novel scientific workflows and system functionalities seek to increase the efficiency, resilience, and pervasiveness of existing systems and applications. Specifically, the proliferation of machine learning/artificial intelligence (ML/AI) workflows, need for processing large-scale datasets produced by instruments at the edge, intensification of near real-time data processing, support for long-term experiment campaigns, and emergence of quantum computing as an adjunct to HPC, have significantly changed the functional and operational requirements of workflow systems. Workflow systems now need to, for example, support data streams from the edge-to-cloud-to-HPC, enable the management of many small-sized files, allow data reduction while ensuring high accuracy, orchestrate distributed services (workflows, instruments, data movement, provenance, publication, etc.) across computing and user facilities, among others. Further, to accelerate science, it is also necessary that these systems implement specifications/standards and APIs for seamless (horizontal and vertical) integration between systems and applications, as well as enable the publication of workflows and their associated products according to the FAIR principles.

97 MATHEMATICS AND COMPUTING↗